Large Language Models

Debug your model. Guardrail its behavior.

Build guardrails and monitors for cyber and CBRN risks at training and inference time, detect reward hacking in training, debug problematic behavior, and intentionally design your models.

What interpretability unlocks

Guardrail

Guardrail your agents

Deploy a multi-staged, low latency guardrail system to prevent unsafe agent behavior, including cybersecurity and CBRN risks, at training and inference time.

Detect

Detect reward hacking
in training

Deploy activation-based and action-based monitors to detect reward hacking. Fix environments, adjust the rubric or reward, or stop runs early.

Debug

Debug model behavior

Trace regressions and unexpected behaviors to specific datapoints, environment bugs, and training runs.

Our research in LLMs

See what your model already knows, and teach it what it's missing.

Research

58% reduction in hallucinations by using features as rewards

We trained Google's Gemma 3 12B using lightweight probes on the model's internal representations as reward signals, cutting hallucinations by as much as the jump from GPT-4o to GPT-5 with no degradation on performance benchmarks.

KEY FINDINGS

58% hallucination reduction with no degradation on performance benchmarks

~90x lower cost than LLM-as-a-judge

Internal representations used directly as training reward signals

Learn more
Research

Using data filtering to mitigate undesired side-effects of post-training

Post-training can introduce undesired side effects that are difficult to detect and even harder to trace to specific training datapoints. We show that a probe-based method can surface concerning behaviors that emerge during LLM post-training, and that probes can identify the datapoints responsible for a specific harmful behavior. Filtering out those datapoints and retraining significantly reduces the behavior.

KEY FINDINGS

Probes trace harmful behaviors back to specific training datapoints

Filtering reduces the harmful behavior by 63% without hurting performance

Identified problematic data sources to omit, leading to 84% reduction in behavior

Learn more

Start researching with Silico

Download Silico for macOS, or talk with us about bringing it to your team and infrastructure.