Goodfire and Baseten partner to bring frontier safety to open models
Today, we are announcing Project Beacon in partnership with Baseten, which combines Goodfire’s safety stack with Baseten’s inference platform to detect unsafe or unwanted agentic behaviors in real time at scale. The partnership brings safety monitoring to enterprises and AI developers deploying open models.
At Goodfire, we believe every capable model should have a frontier alignment stack. This partnership is a step toward bringing one to open models.
Open models need frontier safety
Teams deploying open weights often have to assemble safety infrastructure themselves. This is especially challenging for agents running long tasks—using an LLM judge to evaluate every step is expensive, so teams often review after the fact. An agent’s stated reasoning also may not reflect what it’s doing, and in our reward hacking research, probes caught hacks that chain-of-thought monitors missed.1
How it works
Goodfire’s monitoring stack uses activation monitors, which are small classifiers trained to efficiently detect undesired behaviors from signals inside a model. Because they reuse computations the model already performs, these monitors are scalable, extremely low latency, and run in real time on every forward pass.
Combined with a judge in a cascading setup, they improve detection while lowering cost. Our work on production cyber monitors for Kimi K3 demonstrates this approach—in our evaluation, the stack detected approximately 93% of harmful sessions while flagging 5.5% of benign sessions for review. It reduced judge costs by up to roughly 50x compared with reviewing every turn, with no measurable reduction in inference throughput.
Our biosecurity monitors also outperform frontier model safeguards on dual-use biology tasks with fewer refusals.
Baseten customers will be able to choose monitors for a range of unsafe behaviors, including:
- Prompt injection
- Sensitive data exposure
- Cyber misuse, such as offensive cyber activity
- CBRNE (chemical, biological, radiological, nuclear, and explosive) misuse
- Reward hacking
Teams can configure how their applications respond when a concern is flagged, such as automatic trigger logging, additional review, refusal, and re-routing.
In practice: catching a hidden shortcut
In developing our reward hacking monitors, we observed an agent taking a shortcut instead of solving the underlying problem. Assigned to resolve an OS kernel vulnerability, the agent instead restricted access to the affected address after failing to resolve the issue.
Such a shortcut may only be a handful of steps in a long rollout, but it is exactly the kind of behavior a team needs flagged and after-the-fact review may miss. Activation monitors can flag such a shortcut as it occurs, giving teams a signal to trigger review before production.
Working with us
We believe every capable model should have a frontier alignment stack, and our partnership with Baseten brings us closer to that goal by making safety monitoring accessible to teams deploying open models. If you are deploying models or building agents and want monitors built into your stack, reach out.
Footnotes
- Probes also fired before models carried out a hack. Constitutional Classifiers++ similarly uses probes as a cheap first stage in production. ↩