← BLOG

Goodfire and Baseten partner to bring frontier safety to open models

Today, we are announcing Project Beacon in partnership with Baseten, which combines Goodfire’s safety stack with Baseten’s inference platform to detect unsafe or unwanted agentic behaviors in real time at scale. The partnership brings safety monitoring to enterprises and AI developers deploying open models.

At Goodfire, we believe every capable model should have a frontier alignment stack. This partnership is a step toward bringing one to open models.

Open models need frontier safety

Teams deploying open weights often have to assemble safety infrastructure themselves. This is especially challenging for agents running long tasks—using an LLM judge to evaluate every step is expensive, so teams often review after the fact. An agent’s stated reasoning also may not reflect what it’s doing, and in our reward hacking research, probes caught hacks that chain-of-thought monitors missed.1

How it works

Goodfire’s monitoring stack uses activation monitors, which are small classifiers trained to efficiently detect undesired behaviors from signals inside a model. Because they reuse computations the model already performs, these monitors are scalable, extremely low latency, and run in real time on every forward pass.

Combined with a judge in a cascading setup, they improve detection while lowering cost. Our work on production cyber monitors for Kimi K3 demonstrates this approach—in our evaluation, the stack detected approximately 93% of harmful sessions while flagging 5.5% of benign sessions for review. It reduced judge costs by up to roughly 50x compared with reviewing every turn, with no measurable reduction in inference throughput.

Our biosecurity monitors also outperform frontier model safeguards on dual-use biology tasks with fewer refusals.

Baseten customers will be able to choose monitors for a range of unsafe behaviors, including:

  • Prompt injection
  • Sensitive data exposure
  • Cyber misuse, such as offensive cyber activity
  • CBRNE (chemical, biological, radiological, nuclear, and explosive) misuse
  • Reward hacking

Teams can configure how their applications respond when a concern is flagged, such as automatic trigger logging, additional review, refusal, and re-routing.

In practice: catching a hidden shortcut

In developing our reward hacking monitors, we observed an agent taking a shortcut instead of solving the underlying problem. Assigned to resolve an OS kernel vulnerability, the agent instead restricted access to the affected address after failing to resolve the issue.

Such a shortcut may only be a handful of steps in a long rollout, but it is exactly the kind of behavior a team needs flagged and after-the-fact review may miss. Activation monitors can flag such a shortcut as it occurs, giving teams a signal to trigger review before production.

Working with us

We believe every capable model should have a frontier alignment stack, and our partnership with Baseten brings us closer to that goal by making safety monitoring accessible to teams deploying open models. If you are deploying models or building agents and want monitors built into your stack, reach out.

Footnotes

  1. Probes also fired before models carried out a hack. Constitutional Classifiers++ similarly uses probes as a cheap first stage in production. ↩

Read more from Goodfire

September 30, 2026

We can and must solve alignment

Eric Ho
,
September 28, 2026

A practical guide to sparse autoencoders (SAEs)

No items found.
September 9, 2026

How to build fast, efficient monitors for AI models using probes

No items found.

Research

Training and Deploying Production Cyber Monitors on Kimi K3

October 8, 2026

Better biosecurity monitors for AI agents via protein embeddings

October 1, 2026

Models know when they’re reward hacking — and we can catch them at scale

September 17, 2026
No items found.
Partnerships