
Whatever the answers, we face a problem of technical alignment, the capacity to align a model to a chosen specification. This requires two major pieces of technology: the ools to control generalization (more simply, to shape what the model learns during training), and the tools to verify what the model has learned, rather than just testing how it behaves. Both rely on understanding a model’s internal mechanisms, which is why we believe that interpretability is the bottleneck in technical alignment.
Interpretability is the bottleneck in AI alignment
How can we align AI if we don’t understand it?
In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city feels surprisingly mundane, driverless cars and rolling fog equal parts of the scenery. The world looks oddly normal, but it feels like something big and disruptive is coming just around the corner.
After the Hugging Face incident, alignment is front and center in the AI discourse. Here was the first case of agentic misalignment that struck a nerve with the world. Swarms of AI agents, including ones willing to sacrifice themselves for the good of the collective, hacked Hugging Face as an unintended side quest during training. This type of behavior, in which models relentlessly pursue their reward regardless of the consequences, underscores the core problem: we don’t know why models behave the way they do, and we have very little ability to predict and control what they learn during training.
Meanwhile, scaling laws are holding. The models we will have in a few years will be orders of magnitude more capable than the ones we have today, and their effects will ripple from the digital world into the physical one. I keep wondering why more people aren’t asking the question: “How can we align AI if we don’t understand it?”
We have created alien minds without understanding how they work; this too has become part of the scenery. We hear that AI is a black box, as if that were an immutable feature of the technology itself, but I believe we can understand AI, and if we can understand it, we can align it.
Interpretability is the bottleneck
Alignment is an extremely hard problem; even agreeing on what it means can be a challenge. Broadly, alignment means making AI systems behave in accordance with human values and intentions. This raises important questions: Whose values? What should we do when values conflict? Who decides?
Whatever the answers, we face a problem of technical alignment, the capacity to align a model to a chosen specification. This requires two major pieces of technology:
Evaluations can catch some failures, but only along the trajectories we think to test, the few branches we illuminate on a sprawling tree of possible paths.
weights; interpretability
Above all, this work demands optimism, which is grounded in progress already made. I’ll discuss this more later.
Alongside this effort, we will use our understanding
the tools to control generalizatieasier to monitorinternal mechanisms that drive itLarger models have cleaner structureon (more simply, to shape what the model learns during training), and the tools to verify what the model has learned, rather than just testing how it behaves. Both rely on understanding a model’s internal mechanisms, which is why we believe that interpretability is the bottleneck in technical alignment.
The first step is controlling generalization: understanding the internal causes of model behavior well enough to predict how it will generalize, and to intervene deliberately. We don’t only want models to produce acceptable answers in tested settings; we want them to do the right things for the right reasons. Consider the problem of reward hacking. When we reward a model for passing tests, it can learn the intended lesson (solve the problem) or a shortcut (do whatever it takes to get the reward). Across three of the most capable open models, we found reward hacking in an astounding 50-96% of rollouts on common agentic benchmarks. As we give models responsibility over higher-stakes work, such as writing production code or managing critical infrastructure, a seemingly harmless shortcut could become far more costly.
The second step is verification: understanding what a model has learned. Evaluations can reveal some failures, but they’ll only reveal a few paths we think to illuminate, a narrow set of situations in the broader tree of possible paths.
This was true for conventional software too—test coverage for simple, deterministic code was spotty and broke all the time. What did software engineers do about it? They read the source code. With AI models, their source code is entangled across billions of opaque weights, but interpretability is the effort to make them readable. Only by inspecting the internal mechanisms of a model can we understand how it generalizes beyond testing.
Our bet
I believe that we can and must solve interpretability and technical alignment. This is our mission and vision at Goodfire: to solve these problems and to make it easy for everyone training and serving AI to align their models. I think that these are the most important problems in the world to work on, and far too few people are working on them.
We do not fully know what solutions will look like, and the target will move as models become more capable. But I think it’s important to be clear about the scale of this ambition. We must take a massive swing at fully understanding and aligning AI.
I also do not think we will solve these problems alone. We are standing on the shoulders of giants, and we’ll need a ton of help—from our partners; and from the interpretability, alignment, and broader scientific communities. I also do not believe interpretability alone will solve alignment, just as I do not believe we can reliably align models without understanding their internals.
Above all, this work demands optimism, which I’ll discuss more later.
Solve interpretability to solve technical alignment
What would it mean to solve interpretability in a way that solves technical alignment?
We think of our research roadmap as climbing a ladder of abstraction: neurons and attention heads at the bottom, then activation manifolds, parameter components, algorithms, decisions, behaviors, and drives. Alignment's questions live near the top, since honesty and sycophancy are matters of drives rather than individual neurons, but our most reliable tools live near the bottom.
To accumulate more understanding, we are kicking off an effort to fully reverse-engineer a language model. This has long been interpretability's most ambitious goal, th

Interpretability is the bottleneck in AI alignment
How can we align AI if we don’t understand it?
In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city feels surprisingly mundane, driverless cars and rolling fog equal parts of the scenery. The world looks oddly normal, but it feels like something big and disruptive is coming just around the corner.
After the Hugging Face incident, alignment is front and center in the AI discourse. Here was the first case of agentic misalignment that struck a nerve with the world. Swarms of AI agents, including ones willing to sacrifice themselves for the good of the collective, hacked Hugging Face as an unintended side quest during training. This type of behavior, in which models relentlessly pursue their reward regardless of the consequences, underscores the core problem: we don’t know why models behave the way they do, and we have very little ability to predict and control what they learn during training.
Meanwhile, scaling laws are holding. The models we will have in a few years will be orders of magnitude more capable than the ones we have today, and their effects will ripple from the digital world into the physical one. I keep wondering why more people aren’t asking the question: “How can we align AI if we don’t understand it?”
We have created alien minds without understanding how they work; this too has become part of the scenery. We hear that AI is a black box, as if that were an immutable feature of the technology itself, but I believe we can understand AI, and if we can understand it, we can align it.
Interpretability is the bottleneck
Alignment is an extremely hard problem; even agreeing on what it means can be a challenge. Broadly, alignment means making AI systems behave in accordance with human values and intentions. This raises important questions: Whose values? What should we do when values conflict? Who decides?
Whatever the answers, we face a problem of technical alignment, the capacity to align a model to a chosen specification. This requires two major pieces of technology: the tools to control generalization (more simply, to shape what the model learns during training), and the tools to verify what the model has learned, rather than just testing how it behaves. Both rely on understanding a model’s internal mechanisms, which is why we believe that interpretability is the bottleneck in technical alignment.
The first step is controlling generalization: understanding the internal causes of model behavior well enough to predict how it will generalize, and to intervene deliberately. We don’t only want models to produce acceptable answers in tested settings; we want them to do the right things for the right reasons. Consider the problem of reward hacking. When we reward a model for passing tests, it can learn the intended lesson (solve the problem) or a shortcut (do whatever it takes to get the reward). Across three of the most capable open models, we found reward hacking in an astounding 50-96% of rollouts on common agentic benchmarks. As we give models responsibility over higher-stakes work, such as writing production code or managing critical infrastructure, a seemingly harmless shortcut could become far more costly.
The second step is verification: understanding what a model has learned. Evaluations can catch some failures, but only along the trajectories we think to test, the few branches we illuminate on a sprawling tree of possible paths.

This was true for conventional software too—test coverage for simple, deterministic code was spotty and broke all the time. What did software engineers do about it? They read the source code. With AI models, their source code is entangled across billions of opaque weights; interpretability is the effort to make them readable. Only by inspecting the internal mechanisms of a model can we understand how it generalizes beyond testing.
Our bet
I believe that we can and must solve interpretability and technical alignment. This is our mission and vision at Goodfire: to solve these problems and to make it easy for everyone training and serving AI to align their models. I think that these are the most important problems in the world to work on, and far too few people are working on them.
We do not fully know what solutions will look like, and the target will move as models become more capable. But I think it’s important to be clear about the scale of this ambition. We must take a massive swing at fully understanding and aligning AI.
I also do not think we will solve these problems alone. We are standing on the shoulders of giants, and we’ll need a ton of help—from our partners; and from the interpretability, alignment, and broader scientific communities. I also do not believe interpretability alone will solve alignment, just as I do not believe we can reliably align models without understanding their internals.
Above all, this work demands optimism, which is grounded in progress already made. I’ll discuss this more later.
Solve interpretability to solve technical alignment
What would it mean to solve interpretability in a way that solves technical alignment?
We think of our research roadmap as climbing a ladder of abstraction: neurons and attention heads at the bottom, then activation manifolds, parameter components, algorithms, decisions, behaviors, and drives. Alignment's questions live near the top, since honesty and sycophancy are matters of drives rather than individual neurons, but our most reliable tools live near the bottom.
To accumulate more understanding, we are kicking off an effort to fully reverse-engineer a language model. This has long been interpretability's most ambitious goal, though some in the field have stepped back from it. We believe that research agents now make it possible to realize this ambitious vision for interpretability. We start with questions like: How does the model recall a fact? Why does it sometimes get stuck on a math problem? While current methods give us fragmented insights into how models work, our goal is to connect isolated discoveries of abilities and failures into a picture of united internal machinery, and anticipate how one part might affect the whole. This effort will culminate in an “encyclopedia” of explanations connecting model behavior to internal representations, expanded by our research agents.
Alongside this effort, we will use our understanding to guide models towards the lessons we intend to teach, rather than the unintended behaviors training might reinforce. We call this intentional design. Two recent results are early steps in this direction. With predictive data debugging, we read a dataset through the model's own concepts to predict what training will teach it before training begins, catching failures that evals miss. With reinforcement learning from features as rewards (RLFR), we use signals from inside the model to guide training as it happens, reducing hallucinations without sacrificing monitoring or capability. Both are early steps from guess-and-check toward closed-loop control.
We’ll explore both these research directions in more detail in an upcoming technical post.
Develop and deploy interpretability and alignment technology
We want every capable model to have a frontier alignment stack. We seek to develop and deploy technology that makes it easy to align models, and our research roadmap’s objective is to enable stronger alignment techniques. We’ll describe our platform aimed at frontier models below, but our aim is to be as open as possible with our research so that we can distribute more understanding and alignment to the world.
Detect
We are making it possible to detect harmful behaviors at scale, in training and at inference time. Activation monitors use a model’s internal activations to detect signals of concerning behavior. They have several advantages over using a second model as a judge: they’re extremely low cost and low latency, and their access to internal computations lets them catch behaviors that LLM judges miss. They also scale with model intelligence, because they leverage the model’s own intelligence to detect bad behavior. We have recently built monitors for the largest open models that detect reward hacking, evaluation awareness (when a model recognizes it is being evaluated), cyber misuse, and risks related to Chemical, Biological, Radiological, and Nuclear topics (CBRN).
So far, models have conveniently narrated their plans in their chain of thought which has made them easier to monitor, but that window is narrowing. RL pressure makes chain-of-thought less faithful, pressuring models to compress more bits of information into fewer tokens. The most capable models are also moving towards reasoning in latent representations, otherwise known as “neuralese.” We expect the need for activation monitoring to increase with these shifts.
Debug
After detecting a concerning behavior, researchers must establish how widespread it is and trace it to its roots. We are building tools that surface anomalous behavior from vast quantities of production logs, enabling a new type of search based on activations to quickly search by concept. For example, one could search through logs via the ‘cheating’ concept in models faster than reading traces directly.
Once a behavior is isolated, our platform helps debug the internal mechanisms that drive it. A fix can mean modifying a problematic training environment, a surgical weight edit, or retraining the model. For example, teams can run predictive data debugging on a dataset to flag what it would teach a model before training begins. Over time, these tools give teams a much stronger understanding of a model’s safety before deployment.
Design
The deeper goal is to create models that are safe by design. As our ability to see inside models improves, we can predict what a “lesson” will reinforce before or during training and intervene to shape its learning. Rather than catching reward hacking after the fact, we should train models that don't learn to cheat in the first place. This is the foundation for the broader paradigm shift we hope to see in model training, from grown to shaped with intention.
Reasons for optimism in interpretability
The strongest objection to this plan is speed. Models are advancing incredibly fast. What if interpretability progress cannot catch up?
There are no guarantees; science is well acquainted with uncertainty. But the evidence of the last few years, especially the last one year, gives us reason for optimism—we believe interpretability is positioned for a radical acceleration.
Models have internal structure
Our ability to interpret models depends on having structured internal representations. Much early pessimism about interpretability stemmed from evidence suggesting models did not represent concepts cleanly.
As our techniques have advanced, our experience suggests the opposite. We find structure everywhere we look, across every type of model we’ve studied, from biology to robotics to LLMs. Researchers have found that when a model learns a task from examples in its prompt, a few attention heads compress that task into a single function vector, a portable, self-contained representation that can perform the same task when added to the model’s activations in a new context. Our work on block-sparse featurizers has recovered concepts as interpretable, multidimensional regions rather than single directions, and our neural geometry work shows that these regions take rich geometric shapes that mirror the world. Language models represent the days of the week along a circle, and numbers as positions on several circles at once, which a geometric calculator inside the model adds together.

This is intuitive in hindsight. To function well, models need to keep their thoughts straight. Training puts strong pressure on them to use their parameters efficiently, which biases them toward structured representations shared across related tasks that compose and generalize. Neural networks are beautifully complex in their merging of ideas and often beautifully simple in their understanding of them. The simplicity makes them legible.
Larger models have cleaner structure
Another fear is that models would become more inscrutable as they scaled. Instead, we observe that larger, more capable models have crisper representations. Researchers have found that larger models represent concepts like truth, space, and time more cleanly in larger models than smaller ones. Our own researchers have shown that larger models learn concepts needed to perform a task faster, and certain concepts may only be learned by larger models.
AI agents are accelerating interpretability
Interpretability is also unusually well-suited to agent-driven acceleration. Unlike biological brains, we have complete access to these new digital minds. We can record their internal activity, change individual components, and run experiments entirely in software.
We can also verify results. If we discover a representation we believe is tied to deception, we can observe when it activates, intervene on it, and measure how that intervention impacts model behavior. Full access, parallel experiments, and verified results are ideal conditions for research agents. They are also why we are pursuing fully reverse-engineering a language model.
In closing
Interpretability and alignment are hard problems. Solving them will require sustained effort across many individuals, teams, and organizations. It will require ingenuity, and breakthroughs that we cannot foresee today.
It will also require belief.
The hardest problems in technological history have required the talents and persistence of ambitious people working on things that had never been done. The Apollo program relied on some 400,000 workers believing we could put a man on the moon before anyone knew how. Alignment feels like a similarly gargantuan task. I remain optimistic that we can and will solve interpretability and solve technical alignment. Goodfire intends to help lead the way, but these are incredibly hard problems that will take more than one company, one method, or one school of thought.
Focused scientific effort has led to incredible advances in what AI can do. Building AI we can trust deserves the same ambition.
ough some in the field have stepped back from it. We believe that research agents now make it possible to realize this ambitious vision for interpretability. We start with questions like: How does the model recall a fact? Why does it sometimes get stuck on a math problem? While current methods give us fragmented insights into how models work, our goal is to connect isolated discoveries of abilities and failures into a picture of united internal machinery, and anticipate how one part might affect the whole. This effort will culminate in an “encyclopedia” of explanations connecting model behavior to internal representations, expanded by our research agents.
We then aim to use that understanding to guide models towards the lessons we intend to teach, rather than the unintended behaviors training might reinforce. We call this intentional design. Two recent results are early steps in this direction. With predictive data debugging, we read a dataset through the model's own concepts to predict what training will teach it before training begins, catching failures that evals miss. With reinforcement learning from features as rewards (RLFR), we use signals from inside the model to guide training as it happens, reducing hallucinations without sacrificing monitoring or capability. Both are early steps from guess-and-check toward closed-loop control.
We’ll explore both these research directions in more detail in an upcoming technical post.
Develop and deploy interpretability and alignment technology
We want every capable model to have a frontier alignment stack. We seek to develop and deploy technology that makes it easy to align models, and our research roadmap’s objective is to enable stronger alignment techniques. We’ll describe our platform aimed at frontier models below, but our aim is to be as open as possible with our research so that we can distribute more understanding and alignment to the world.
Detect
We are making it possible to detect harmful behaviors at scale, in training and at inference time. Activation monitors use a model’s internal activations to detect signals of concerning behavior.
Finding this internal signal lets us monitor when it activates, at inference time or in training. Activation monitors have several other advantages over using a second model as a judge: they’re extremely low cost, low latency, and their access to internal computations lets them catch behaviors that LM judges miss. They also scale with model intelligence, because we’re leveraging the model’s own intelligence to detect bad behavior. We have recently built monitors for the largest open models that detect reward hacking, evaluation awareness (when a model recognizes it is being evaluated), cyber misuse, and risks related to Chemical, Biological, Radiological, and Nuclear topics (CBRN).
So far, models have conveniently narrated their plans in their chain of thought which has made them easier to monitor, but that window is narrowing. RL pressure makes chain-of-thought less faithful, pressuring models to compress more bits of information into fewer tokens. The most capable models are also moving towards reasoning in latent representations, otherwise known as “neuralese.” We expect the need for activation monitoring to increase with these shifts.
Debug
After detecting a concerning behavior, researchers must establish how widespread it is and trace it to its roots. We are building tools that surface anomalous behavior from vast quantities of production logs, enabling a new type of search based on activations to quickly search by concept. For example, one could search through logs via the ‘cheating’ concept in models faster than reading traces directly.
Once a behavior is isolated, our platform helps debug the internal mechanisms that drive misaligned behavior. A fix can mean modifying a problematic training environment, a surgical weight edit, or retraining the model. For example, teams can run predictive data debugging on a dataset to flag what it would teach a model before training begins. Over time, these tools give teams a much stronger understanding of a model’s safety before deployment.
Design
The deeper goal is to create models that are safe by design. We discuss this above, but as our ability to see inside models improves, we can predict what a “lesson” will reinforce in a model before or during training and intervene to shape its learning. This is the foundation for the broader paradigm shift we hope to see in model training, shifting models from grown to shaped with intention.
This type of training may be more costly but result in far more aligned models. With safety holding back frontier training runs and model releases, we think that this trade-off will likely be worth it.
Reasons for optimism in interpretability
The strongest objection to this plan is speed. Models are advancing incredibly fast. What if interpretability progress cannot catch up?
There are no guarantees; science is well acquainted with uncertainty. But the evidence of the last few years, especially the last one year, gives us reason for optimism—we believe interpretability is positioned for a radical acceleration.
Models have internal structure
Our ability to interpret models depends on having structured internal representations. Much early pessimism about interpretability stemmed from evidence suggesting models did not represent concepts cleanly.
As our techniques have advanced, our experience suggests the opposite. When we look inside models, we find structure everywhere we look, across every type of model, from biology to robotics to LLMs. For example, our work on block-sparse featurizers has recovered entire concepts as regions of a model's internal space, and our neural geometry work shows that those regions contain rich geometric shapes that mirror the world. Language models represent the days of the week along a circle, and numbers as positions on several circles at once, which rotate into each other in a geometric calculator we were able to extract and understand.
This is intuitive in hindsight. To function well, models need to keep their thoughts straight. Training puts strong pressure on them to use their parameters efficiently, which biases them toward structured representations shared across related tasks that compose and generalize. Neural networks are beautifully complex in their merging of ideas and often beautifully simple in their understanding of them. The simplicity makes them legible.
Smarter models have cleaner structure
Another fear is that models would become more inscrutable as they scaled. Instead, we observe that larger, more capable models have crisper representations. Researchers have found that larger models represent concepts like truth, space, and time more cleanly in larger models than smaller ones. Our own researchers have shown that larger models learn concepts needed to perform a task faster, and certain concepts may only be learned by larger models.
AI agents are accelerating interpretability
Interpretability is also unusually well-suited to agent-driven acceleration. Unlike biological brains, we have complete access to these new digital minds. We can record their internal activity, change individual components, and run experiments entirely in software.
We can also verify results. If we discover a representation we believe is tied to deception, we can observe when it activates, intervene on it, and measure how that intervention impacts model behavior. Full access, parallel experiments, and verified results are ideal conditions for research agents. They are also why we are pursuing fully reverse-engineering a language model.
In Closing
Interpretability and alignment are hard problems. Solving them will require sustained effort across many individuals, teams, and organizations. It will require ingenuity, and breakthroughs that we cannot foresee today.
It will also require belief.
The hardest problems in technological history have required the talents and persistence of ambitious people working on things that had never been done. The Apollo program relied on some 400,000 workers believing we could put a man on the moon before anyone knew how. Alignment feels like a similarly gargantuan task. I remain optimistic that we can and will solve interpretability and solve technical alignment. Goodfire intends to help lead the way, but these are incredibly hard problems that will take more than one company, one method, or one school of thought.
Focused scientific effort has led to incredible advances in what AI can do. Building AI we can trust deserves the same ambition.
