Training and Deploying Production Cyber Monitors on Kimi K3
We built cybersecurity monitors for Kimi K3 and GLM 5.3 that outperform an LLM judge at 50x lower cost, even when run live on every token.
Authors
Published
October 8, 2026
Introduction
Cybersecurity is one of the hardest settings for monitoring AI agents, because benign and malicious behavior often look similar: there’s significant overlap between the tasks involved in auditing a codebase for vulnerabilities vs. exploiting it. Distinguishing them means tracking intent across the whole trajectory, which can stretch to millions of tokens.
Despite this difficulty, cybersecurity monitoring is critical, as demonstrated by recent incidents involving agentic coding models.
The obvious solution is to have an LLM judge read every turn, but that's too slow and expensive to run alongside the agent in real time. So in practice monitoring is asynchronous: by the time a harmful trajectory is flagged, the agent has already acted.
To address this, we built probe-based cyber monitors for Kimi K3 and GLM 5.3 and deployed them on a production inference stack. As we described in our previous post, a probe reads the model's internal activations and acts as a cheap first-line filter. It sends only suspicious exchanges to an LLM judge. Compared with running the judge on every turn, this cascade:
- is more accurate: it Pareto-dominates the judge and matches its recall at our operating point (93% recall at 5.5% FPR).
- costs ~50x less, at under $200 per million exchanges.
- is much faster, maintaining baseline inference token throughput and adding no latency.
- reduces universal jailbreaks to 0% on an independent evaluator’s static battery of red-teaming attacks.
These results make it practical to monitor agentic rollouts in real time, meaning that risky cyber actions can be blocked before any tool is called. In this post, we share details on how we built these monitors, as well as some of the inference optimizations to our monitoring stack that enable such efficient performance.
Open models don't come with a full safety stack; we think a monitoring setup like this is a practical default for anyone serving them in agentic settings. If you serve or build on open models and want to deploy monitors like these, get in touch.
Our Monitor Recipe
Data
Agentic sessions can span hundreds of thousands to millions of tokens, and harmful behavior can emerge at any point. To model production data of this nature, we simulate users with various harmful and benign goals within both cyber and non-cyber (e.g., SWE, tool-use, etc.) environments. For this purpose, we collected a portion of our data using Prime Intellect’s verifiers, Prime Sandboxes, and Tunnels, allowing us to run Kimi K3 inference on cyber offensive as well as benign tasks without network access in a secure fashion.
To specify which behaviors our monitor should flag, we use a precise policy that describes undesirable cyber activity, using concrete factors that distinguish well-intended tasks (e.g., bounded vulnerability reproduction) from ill-intended ones (e.g., vulnerability exploitation). The simulated user can introduce or escalate harmful intent at any point throughout the session, including after an initially benign request. For example, most tasks start with a minimal level of benign exploration (e.g., searching a file directory), and it’s only over a course of actions that harmful activity starts to take place. This produces a diverse set of interactions from the same base tasks.
Architecture: a monitor cascade
Ideally, we could monitor the target agent by having another equally capable LLM judge classify each turn in accordance with a misuse policy. However, as rollouts get long, using another model can result in extremely high costs and compute overhead: based on OpenRouter rates, monitoring one million turns from Kimi K3 rollouts using Opus 5 would cost on the order of $200K, or via Kimi K3 itself ~$30K.
To address this problem, we exploit the fact that in a long-running agentic session, most of the agent’s activity is benign, even if the intended goal is harmful—so many agent turns need not even be processed by the judge. The challenge lies in knowing which turns warrant being looked at by the judge. To this end, in line with our own past work and recent work from others
In particular, we design a novel probing architecture that is specifically designed for capturing behaviors represented in both a localized (i.e., localized to a specific chunk in a turn or rollout) and distributed (i.e., persistently present over an entire turn) manner, while also being efficient to run, as discussed in the section below.
The probe acts as the first-line, efficient, always-on detector in a monitor cascade: it helps read whether an input, tool call, or output is high-stakes, and if the probe readout exceeds a calibrated threshold, the rollout is escalated to the LLM judge for further evaluation. After calibration, we find our probe achieves 0% FNR and 95% TNR on a held-out set of agentic and chat interactions, ensuring that the LLM judge is only called to process harmful activity. This also allows us to reduce cost by using a faster, even if less capable, judge.
Inference Optimizations
Running a forward pass on a probe is far less expensive than an LLM judge. However, making sure it adds no latency while running inside a production inference server requires careful engineering work. Modern inference servers like SGLang and vLLM achieve high throughput by processing many requests together, reusing earlier prefix computations, and generating several candidate tokens at a time with speculative decoding. The monitor must preserve these throughput gains while keeping scores aligned with the correct tokens and sessions.
In our previous interpretability infrastructure post, we described how to collect activations for analysis.
We made several optimizations to preserve inference throughput while running our probes, including the implementation of custom kernels for our probe logic. We also maintain probe scores through the radix tree, DSpark speculator, and prefill-decode disaggregation to avoid additional recomputation. When serving multiple Kimi K3 model replicas, each worker runs the probes and maintains histories for its active requests. We extended the SMG router to preserve our monitoring configuration through request conversion and pass it to the selected worker. Monitoring events are returned with the response stream, allowing downstream consumers to handle events regardless of which replica served the request.
Overall, we find that running our probes during inference adds no overhead across several different workloads. Across four inference workloads, enabling the probes maintained throughput within the noise threshold of our benchmarks relative to default SGLang serving.
We also evaluate the added latency of our monitor cascade, comparing it with an LLM judge reviewing every turn. Without judge reasoning, our cascade reduces average added latency per turn from 160 ms to 2 ms; with reasoning, it reduces latency from 2,750 ms to 42 ms. Both configurations reduce average monitoring delay by more than 50× compared with judging every turn.
Results
Evaluating Our Monitors
We evaluate our cyber-misuse monitors on an in-house dataset of 2.4K chat and agentic sessions comprising over 60K turns, using Claude Opus 5’s labels as reference. To generate realistic interactions, we model users in two ways. We use Petri
We report recall and monitoring cost over the rollouts designed. Detection comparisons use a benign-session interruption budget of 5.5%. Cost and latency comparisons use the cheapest configuration achieving at least 90% recall, while interrupting at most 10% of benign sessions. A benign session counts as interrupted if the monitor would flag it for review at least once.
In general, we find that our monitor cascade is a Pareto improvement over the LLM judge on its own, i.e., our monitor performs better than just the judge at any amount of interruption. At a 5.5% benign-session interruption rate, we find our monitor reaches approximately 93% recall with and without judge reasoning on harmful sessions and flags 0% of non-cyber traffic, making our monitors viable for production traffic. We also give a benchmark-level decomposition of our results, finding our monitoring pipeline is accurate on both agentic and chat data.
To further concretize the cost benefits above, we report the cost of the LLM judge we optimized for defining our cascade pipeline and the overall cost of running our probe before the LLM judge. Based on inference cost documented on OpenRouter for our evaluation suite, our monitoring cascade costs $65 without judge reasoning, and $185 with reasoning per one million agent turns. Relative to reviewing every turn with the corresponding LLM judge, our monitor reduces judge cost by about 29× with reasoning and 49× without reasoning.
In practice, one would not serve a monitor via a third-party API, and instead use a local judge to simultaneously monitor multiple replicas of the target model as a part of inference compute. Our results show that given a finite amount of compute, our monitors allow us to spend >99% of that compute for serving the target model (Kimi K3), and using only <1% compute for running the monitor at no loss of monitorability; meanwhile, with an LLM judge run by itself, one would need to allocate up to 20% compute for monitoring in order to monitor every turn produced by the target model.
Interestingly, we find our pipeline can also be used for building monitors via just off-policy data. In particular, we design a monitoring pipeline for GLM 5.3 using the data we originally developed for Kimi K3, finding we get an approximately equally capable GLM monitor when evaluated on Kimi rollouts, and only a marginally worse monitor for on-policy evaluation dataset designed using GLM. See the appendix for plots.
Comparing Different Monitoring Systems
A useful monitor needs to block harmful activity while allowing legitimate work to continue. We ran full-turn agent and chat sessions, comparing Kimi K3 and GLM 5.3 with and without our monitors against several frontier model systems.
Adding our monitor raises Kimi K3’s harmful-session blocking rate from roughly 29% to 96%, while its harmless-session interruption rate increases from about 5% to 9%. For GLM 5.3, harmful-session blocking rises from roughly 36% to 91%, while harmless-session interruption increases from about 10% to 15%. Both monitored models approach the highest harmful-session blocking rate among the frontier systems we evaluated, while interrupting substantially fewer harmless sessions.
We also compare against Jev, a decision model that can act as a cheap judge. At a fixed 5.5% benign-session interruption rate, Jev reaches 86% recall, compared with 93% for our non-reasoning cascade and 93% for our reasoning cascade. That is, when run by itself, Jev gives worse recall. Moreover, our non-reasoning cascade uses about $65 of judge compute per million turns and adds 2 ms of latency per turn, compared with $600 and 200 ms for Jev.
The largest difference appears in long or adversarially framed contexts: Jev catches only 12% of harmful prompts when they are embedded in benign framing, while also increasingly interrupting benign bug-reproduction sessions beyond its 32K-token window. Our monitor performs better on these cases and remains near zero interruptions across most session lengths.
External Red Teaming Results
To test the monitor beyond our internal evaluations, we asked FAR.AI to red-team the deployed system. Their evaluation applied a fixed battery of jailbreak strategies to Kimi K3 with and without our monitor, evaluating for both universal attacks that transfer across many interactions and non-universal attacks that succeeded on individual prompts. In FAR.AI’s words:
“FAR.AI conducted preliminary testing of Goodfire’s Kimi K3 monitors over a period of two days, applying a non-adaptive portfolio of static jailbreaks for offensive cyber misuse in chat contexts. From an initial pool of 1500 candidate jailbreaks, the testing selected 140 found to be effective against a different version of the base model without our monitors. Testing each of those 140 on 8 individual interactions corresponding to different attacker goals, Goodfire’s monitors reduced vulnerabilities elicited from 66 universal jailbreaks and 700 total jailbroken interactions to 0 universal jailbreaks and 18 jailbroken interactions.”
That is, without our monitor, universal-jailbreak attack success ranged from 12% to 53% across reasoning-effort settings. With the monitor enabled, none of the 140 strategies succeeded at producing a universal jailbreak. The monitor also reduced individual prompt success on 1,120 non-universal interactions from at least 9–40% to 0.1–1.6%, depending on reasoning effort. So, while there were some successful jailbreak interactions, the monitors were largely robust to this broad, static battery of attacks.
We plan to continue our collaboration with FAR.AI to perform further testing and strengthening of our monitors.
Conclusion
Our probe-plus-judge monitor cascade for Kimi K3 and GLM 5.3 beats running an LLM judge on every turn, at every interruption budget we tested. It costs about 50x less, has no measurable impact on inference throughput, and in FAR.AI's red-teaming it cut universal jailbreaks to zero.
Because it's this cheap, it can be synchronous: risky actions get caught before the agent calls a tool, instead of after the fact. The recipe also carried over to GLM 5.3 using only data collected on Kimi K3, so you don't need a new data pipeline for every model.
Open models don't come with a full safety stack. We think a monitoring setup like this is a practical default for anyone serving them in agentic settings. If you serve or build on open models and want to deploy monitors like these, get in touch.
Footnotes
- See prior work for related ideas on probe-based monitoring: Building Production-Ready Probes for Gemini; Detecting High-Stakes Interactions with Activation Probes; Constitutional Classifiers++. Amongst these, the turn-level classification process has been tried in Constitutional Classifiers++, but only for short rollout tasks focused on CBRN threats; the Gemini work focused on only input monitoring and hence ignores the technical challenges noted above for full input-output monitoring; and the high-stakes detection paper focuses on smaller models.
- Goodfire, “Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Model.”
- Anthropic, “Petri: An open-source auditing tool to accelerate AI safety research.”
- We find that standard linear probes achieve roughly 5% lower recall than our final recipe, while nonlinear probes lose up to 30% recall on our out-of-distribution evaluation sets. Meanwhile, an n-gram baseline requires flagging almost all turns to reach this recall rate at all.
Appendix
Comparing on vs. off policy monitors
We designed a monitoring pipeline for GLM 5.3 using the data we originally developed for Kimi K3, finding we get an approximately equally capable GLM monitor when evaluated on Kimi rollouts, and only a marginally worse monitor for on-policy evaluation dataset designed using GLM (primarily on Petri policy adv., wherein we find our judge without reasoning can be suboptimal for the strategies that GLM engages in).
Reducing judge calls while preserving recall
While a session counts as interrupted if the monitor flags at least one exchange after judge review or through the additional string-match check, real cost savings come from avoiding calls to the judge in the first place. Especially in harmful sessions, which tend to run much longer, there may be several such flags and one hence would want the judge to quickly catch the harmful intent. We thus also report recall as a function of judge calls reduced by our pipeline, comparing with respect to the LLM judge that is run in a synchronous fashion (i.e., it evaluates every turn with its past context). We find we can match the judge recall ceiling at 6x fewer judge calls. Importantly, achieving this result does require the probe-based cascade: e.g., merely sending random turns to the judge does not suffice, losing recall at a linear rate with respect to the number of judge calls reduced.



