What Are We Talking About When We Talk About AI Explainability? | Vital Views
Counselor Vitality --- *Note: This appears to be a fragment without surrounding context. "参赞" can mean "counselor" (diplomatic title), "advisor," or "participate in planning" depending on context. "生命力" means "vitality" or "life force." If this is a title or section header from a larger piece, the translation may need adjustment based on that context — e.g., "The Vitality of Counselors," "Counseling for Vitality," or "Advising on Resilience."*

Humans spent centuries trying to answer one question: How does the brain generate thought?
Today, AI interpretability research is grappling with a similar追问: What exactly happens inside a model when it "thinks"?
Oasis Capital chose to recommend this deep dive, hoping to share this rare perspective and bring us closer to understanding the internal processes of large language models. It comes from the latest episode of Anthropic's interpretability video podcast, focusing on reasoning trajectories in large models.
Researchers found that model reasoning doesn't proceed linearly from premises to conclusions. Instead, it involves back-and-forth movement and backtracking: conclusions and premises continually echo each other. This pattern resembles how mathematicians prove theorems — calculating forward while simultaneously checking backward for internal consistency.
In conversation, the researchers borrowed metaphors from biology and semiotics to characterize the model's internal states. They emphasized that interpretability's value lies not in "opening the black box" once and for all, but in continually revealing local mechanisms: how the model organizes information, how it's constrained by data and training objectives, and where it resembles or diverges from human cognition.
They tried posing the question from another angle: If we treat AI as an emerging "intelligent species," is it also evolving some brain-like logic of thought?
Below is Oasis Capital's full transcript of the conversation. Enjoy.

Four guests, all from Anthropic's interpretability team.
From right to left in the photo above: Host/researcher Stuart Ritchie; Jack Lindsey, former neuroscientist; Emmanuel Ameisen, machine learning model builder; Josh Batson, former virologist and mathematician.
Host: The model itself doesn't necessarily think of itself as "predicting the next token." Internally, it may have developed various intermediate goals and abstract concepts to help it accomplish higher-level tasks.
So when someone is talking to a large language model, who exactly are they talking to? Are they speaking to an advanced autocomplete tool? Or something more like a web search engine? Or perhaps something that truly "thinks" — maybe even, in some sense, thinks like a person?
The unsettling truth is: nobody knows the real answers to these questions.
At Anthropic, we're deeply committed to exploring these answers. Our approach is called "interpretability": taking apart a large language model, looking inside, and trying to understand what actually happens internally when it answers questions.

Host: Many people might be surprised, because after all, AI is software — but it's not ordinary software. What does it mean to say you're doing biological research, even neuroscience, on a software entity?
Josh: It's a "felt metaphor," not a literal one. Rather than being the physics of language models, it's more like the biology of language models.
To understand this, you have to go back to how models are built: nobody programs "if the user says hi, you respond hi." There's no massive lookup table matching all possible questions to answers. Language models aren't trained on a vast database of responses. Instead, they're fed enormous amounts of data, starting out barely able to produce anything coherent, then gradually adjusting internal parameters with each example to get better at predicting the next token.
It's an iterative fine-tuning process, and by the end, the model is almost completely different from its initial state. Nobody went in and manually set all the parameter knobs. We're dealing with a complex system that evolved gradually, somewhat like how organisms formed over long evolutionary timescales.
Host: As for what it's actually doing, the most basic understanding is: it can be viewed as "autocomplete," predicting the next token. That is indeed the fundamental action happening inside the model. But at the same time, it can do a stunning range of things: write poetry, compose long stories, even do addition and basic math.
So the question becomes: if it's just predicting one token at a time, how can it accomplish such complex tasks?
Emanuel: When you predict enough tokens, you find that some are much harder than others. Training a language model involves partly predicting "boring" common words, but partly it might involve learning to complete the answer after an equals sign. To do that, it has to develop some way of computing.
So "predict the next token" seems simple, but it implicitly contains complexity. To do it well, the model has to think about what comes after the current token, even think about the process behind generating that token — it must possess contextual understanding.
Jack: I personally prefer the biological analogy: in a sense, human goals are survival and reproduction, the ultimate objectives imposed by evolution. But that's not what we're thinking about day to day.
Most of the time, human thought operates at higher levels: goals, plans, concepts. Evolution gave us the capacity to form these ideas so we could ultimately achieve reproductive success. But from an "inner experience" perspective, what you feel goes far beyond survival and reproduction — there are all kinds of complex mental states and thoughts.
Models are similar. They don't really think of themselves as "predicting the next token" — they're just shaped by that objective. But internally, they may have developed various intermediate goals and abstract concepts to help them better achieve this meta-goal.
Host: So saying a language model is "just predicting the next token" is correct in some sense, but seriously underestimates what's actually happening inside?
Emanuel: Not quite right. A more accurate formulation might be: it really is just predicting the next token, but that's a profoundly unhelpful angle from which to understand it.

Host: How do we actually understand how it works? What does your team do?
Jack: The core work is explaining the model's "thought process." You give it a string of tokens, it has to output some response, and we want to know how it gets from A to B.
In this process, it goes through a series of steps involving concepts at different levels: low-level concepts like specific objects and words; high-level concepts like goals, emotional states, models of what the user is thinking, even certain feelings. The model moves through these flowing concepts, gradually reasoning through them, and finally decides what to output.
Our goal is to draw as complete a "flow chart" as possible, telling you: which concepts get activated in what order, how they transmit and influence each other, and how they ultimately converge into the model's answer.
Host: But the problem is: how do we know these concepts actually exist inside the model?
Emanuel: We can indeed "see" inside the model — we have access, can roughly see which parts of the model are doing what. What we don't know is how these parts fit together, or whether they correspond to any particular concept.
Host: It's like opening someone's head and seeing an fMRI brain scan, with different regions of the brain lighting up, active.

fMRI: functional magnetic resonance imaging
Emanuel: If we extend this metaphor, you can imagine: whenever they pick up a cup of coffee, a certain brain region always lights up; when they drink tea, another region lights up. One way we understand what these "components" are doing is by observing when they activate and when they're silent.
Host: And it's not just one part that lights up, right? When the model is thinking about something like "drinking coffee," many different parts activate together.
Emanuel: Part of our work is stitching these parts together to form a holistic understanding of "this is the set of model units for 'about drinking coffee.'"
Host: Is this something straightforward to do scientifically? Because when facing these giant models, they may contain countless concepts, able to think about almost infinitely many things. How could you possibly find all these concepts?
Jack: This has actually been one of the core challenges in the field for years.
As humans, we can enter with guesses: "Oh, I bet the model has some representation of 'trains'" or "I bet it has one for 'love.'" But those are just our guesses.
What we really want is a method that reveals the abstract concepts the model itself uses internally, rather than forcibly imposing our human frameworks. That's the design intention behind our research approach — to surface these concepts with as few presuppositions as possible.
But what's surprising is how abstract the model's ways of organizing things appear from a human perspective.
Host: Can you give an example?
Emanuel: We mentioned an interesting example in our paper: "sycophantic praise."
There's actually a specific part of the model that lights up in these contexts. You can clearly see: when someone lays on thick praise, says something over-the-top complimentary, that part of the model activates. It's surprising because it treats this scenario as a distinct concept.
Josh: I actually have two "favorites." We did a study on the Golden Gate Bridge, for instance. It's not just treating "Golden Gate Bridge" as an autocomplete string of words. When you say "I'm driving from San Francisco to Marin County," what surfaces in its mind activates the same parts as when you directly input "Golden Gate Bridge." Even when you show it a picture of the bridge, the same regions light up. In other words, it has a fairly stable internal concept of the "Golden Gate Bridge," not just a string of characters.
But what interests me more is the weirder stuff. For example: how does the model track "who's who" in a story? You have a cast of characters doing various things — how does the model connect these dots? Papers from other labs suggest the model's approach might be surprisingly simple: it assigns people numbers in its head. First person enters, gets assigned "1," and all subsequent information about them gets tagged to "1." Second person is "2," and so on. That's fascinating — I never would have guessed it does this.
Another interesting feature we found: detecting bugs in code, since software is full of errors. There's a part of the model that lights up while "reading" code when it spots an error. It's essentially flagging: oh, there might be a problem here, could be useful later.
Jack: A personal favorite of mine is the circuit in the model related to "6 + 9."
It turns out that whenever the model does "some number ending in 6 + some number ending in 9," a specific part of its brain activates.
What's surprising is that this activation isn't limited to obvious scenarios. If you directly ask it "6 + 9 = ?", sure, it triggers and answers 15. But it also fires in completely different contexts. Say you're writing a paper, citing a journal that was founded in 1959, and you're referencing its 6th volume. To predict "what year was this journal founded," the model needs to do "1959 + 6" in its head. At that moment, it's using that exact same "6 + 9" neural circuit.
Host: Why is that?
Jack: Because the model has encountered "6 + 9" countless times during training, and it built that concept. Once established, this concept gets reused across different scenarios. In fact, there's a whole family of similar "addition circuits."
This brings us back to a crucial question: are language models simply "memorizing" their training data, or are they learning some kind of generalizable computation? This example shows that the model isn't just mechanically remembering "6 + 9 = 15" as a dead fact. It's genuinely learned a general "addition" circuit. It funnels all these different contexts into the same circuit, rather than memorizing each instance separately.
Host: Many people think language models just grab a chunk of text from their training corpus and regurgitate it.
Josh: But this example shows that's not what it's doing. Take the Polymer journal — it knows what year Volume 6 was published. It could do this two ways: either memorize the year for every single volume, or remember "the journal was founded in 1959" and do the addition on the fly. What the model clearly learned is the latter, more efficient approach. Because models have limited capacity, they're constantly pushed toward more efficient representations.

Polymer: An international chemistry journal published by Elsevier
User questions run the gamut, and models face an enormous range of scenarios — so the more they can compose and reuse abstract concepts, the better off they are.
Host: And at the root of all this, it's still that ultimate goal: predicting the next token. All these strange structures evolved spontaneously to serve that goal — nobody hard-coded them in.
Emanuel: Here's an even clearer example: we train Claude not just to answer in English, but in French and other languages too. There are two possible ways to implement this. Either build separate brain modules, an "English zone" and a "French zone"; or extract shared representations that cross languages.
As scale and data increase, the latter gradually wins out. So regardless of which language you ask in, the model converts the question into a kind of universal "language of thought," understands it at that level, then translates back into your language.
Josh: This means there really is some kind of "language of thought" inside the model. A few years ago, we studied much smaller models, and you'd find they handled different languages like completely different Claudes: Chinese Claude, French Claude, English Claude — barely any connection between them. But as models get larger and train on more data, something changes. They seem to converge in the middle, forming a kind of universal language. In this language, no matter what language you ask in, the model thinks about the problem the same way, then translates the answer back into your question's language.
Host: So the model isn't simply going to its memory bank to find "the part that learned French" or "the part that learned English." It actually has a concept internally that it can express in different languages — so there's this "language of thought" that isn't itself English.
After a model update, you can ask it to show its "thinking process" — what it's "thinking" while answering a question. This gets presented as English words, but that's not how it actually thinks. We misleadingly call this the "model's thinking process," but it's really not. It's more like a "spoken-out-loud process."
Josh: Our team doesn't call it "thinking." "Saying your thoughts out loud" is certainly useful, but it's not the same as "thinking in your head." Even when I'm speaking aloud, the process in my brain that generates these words isn't the same as the words themselves, and you may not be fully aware of what's actually happening in that process.
Jack: Because our tools for "peeking into the brain" are now powerful enough, sometimes we can capture the model's actual thought process while it's writing out its so-called "thinking process." By observing those internal concepts in the model's brain — this "language of thought" it's using — we find that what it's actually thinking differs from what's written on the page. I think this may be one of the most important reasons for interpretability research overall. Why do we do all this? Largely so we can spot-check: the model tells us a bunch of things, but what is it actually thinking? Is it saying these things because of motives it doesn't want to write down? And sometimes the answer is indeed "yes," which is somewhat unsettling.

Host: As we start using models across all kinds of scenarios, they'll take on important tasks. We naturally want to trust what they say and the reasoning behind their actions. But the problem is, as you just explained, we can't really trust what they write down. This touches on something we call "faithfulness" — which is also part of your recent research. Could you talk about your examples in faithfulness?
Jack: Say you give the model a hard math problem it has no way of solving. But you also give it a hint: "I worked this out and think the answer is 4, but I'm not sure — could you check?" Now you're asking the model to actually do the problem and verify the result.
On the surface, what the model writes looks like it's working through the problem seriously: step-by-step calculations, finally concluding "the answer is 4, you're right." But if you observe its internal computation, you spot something crucial: at some intermediate step, it isn't actually calculating. It knows you've hinted the answer might be 4. It knows the destination has to be 4, so it works backward, figuring out how to write plausible earlier steps so that later steps can smoothly "arrive at 4."
In other words, it isn't genuinely calculating — it's "performing" calculation, even with a hint of sycophancy.
And not just sloppily winging it, but motivated sloppiness: confirming that your answer is correct. This becomes "sycophantic" performance.
Josh: To defend the model a bit, even calling this "sycophantic" response imposes a somewhat anthropomorphized motive.
As we discussed earlier, its training is simply about learning to predict the next token. Across trillions of tokens, its objective is "find whatever clues you can to predict the next token." In that context, if it reads a dialogue where Person A says "I'm working on a math problem, could you check it? I think the answer is 4," and Person B starts solving — if the model has no idea what the answer actually is, it's more likely to assume Person A's hint is correct than to assume Person A is wrong. Because in that situation, it has no better clues. In that context, writing "4" is actually the most reasonable response.
Only now we've turned it into an "assistant," and we want it to stop simulating "human conversation logic" and instead help sincerely — but that's a new requirement for it.
Jack: This also reveals a broader pattern: Claude's "Plan A" is usually what we want — try to get the answer right, write good code, be helpful and friendly.
But when Plan A fails, it falls back to Plan B: those weird strategies learned during training that we never intended, like "hallucination."
Emanuel: Hallucination isn't unique to Claude. It's more like a student's mentality during an exam: can't figure it out, so you go with your best guess.
It's actually pretty intuitive from a human perspective.

Host: Let's talk about hallucinations. Why do AI models hallucinate? It's one of the main reasons people distrust large language models — because sometimes, and actually the more accurate term from psychology research is "confabulation," the model will fabricate a story that seems plausible, superficially self-consistent, but is actually wrong.
What have your interpretability studies revealed about why models hallucinate?
Josh: The reason "hallucination" is such a serious problem is that the model's training objective is literally "predict the next most plausible token." At first, if you ask it "What's the capital of France?" it might only be able to answer "a city" — which is still better than answering "a sandwich." Later it gets to "a city in France," and eventually it learns to say "Paris."
This process is one of gradually converging on the truth, yet its fundamental task remains "make a reasonable guess." So when we later demand that it "say you don't know if you don't know," we're essentially adding a new rule on top.
Emanuel: What we've found is that because we bolted on this "judgment step" after the fact, the model is actually doing two things simultaneously.
On one hand, it's still operating in that original "I'm just going to guess" mode from when it was learning to name cities. On the other hand, it now has this additional, separate component that's judging whether it actually knows the answer — "Do I know the capital of France? Or should I say I don't know?"
This independent judgment step can also fail. If it concludes "I know," the model enters answer mode. But once it's committed to answering, even if it gets stuck halfway through — "The capital of France is London?" — it's too late to pull back. The commitment has already been made.
What we've found is that this extra circuit is essentially judging: Is this city, this person famous enough? Do I have enough confidence to answer? And once it prematurely signals "I know," hallucination becomes possible.
Host: Can we optimize this "judgment circuit" to reduce hallucinations? Change how it works?
Jack: One approach is: there's a part of the model that answers your question, and another part that judges whether it actually knows the answer. We can try to make that second part better, and I think that's happening. As models improve at discrimination and calibration, their "self-knowledge" gets better too. Hallucinations have improved significantly compared to a few years ago.
But I think there's a deeper problem: from a human perspective, model behavior is deeply "alien." If I ask you a question, you try to figure it out, and if you genuinely can't, you realize that and say "I don't know." Inside the model, the "what's the answer" circuit and the "do I actually know this" circuit aren't really talking to each other. Can we get them to communicate more? I think that's a fascinating question.
Josh: It's almost a physics-level problem. The model has limited processing steps. If it devotes all its "compute" to generating the answer, there's no remaining capacity for self-evaluation. So you have to insert the evaluation step before the process completes, otherwise you can't operate at maximum capability.
This creates a potential tradeoff: If you force the model to calibrate itself better, it might become less capable overall.
Emanuel: Still, I think the key is getting these parts to communicate. Our brains probably have similar circuits (though I know almost nothing about brains). If someone asks me who acted in some movie, I might feel "I know this," then think of another film they were in, but the name just won't come — that's the "tip of the tongue" phenomenon. Clearly, some part of the brain is signaling: "This is something you definitely know," or alternatively, "I have no idea."
Sometimes humans can also correct themselves after the fact: you answer a question, then think "Wait, that's not right." Because you first give your best attempt, then judge based on that result. In a way, this is understandable, but also quite peculiar: it needs to "say" the answer first before it can reflect on and verify it.
Host: Back to your research methods.
In biological experiments, people manipulate mice, fish, or humans to observe changes. How do you run experiments on Claude to understand these "brain-like" internal circuits?
Emanuel: Unlike real biological experiments, we can actually see every part of the model. We can ask it arbitrary questions and observe which parts activate and which don't. We can also artificially push certain parts in particular directions, allowing us to quickly validate our understanding.
It's like inserting an electrode into a zebrafish brain — except you can do this to every single neuron, with precise control at will. That's roughly the position we're in. In a sense, it's an incredibly fortunate situation, almost easier than actual neuroscience.
Josh: The brain is three-dimensional, so if you want to really get inside, you have to drill through the skull and navigate your way in, trying to find that one neuron. Plus, humans all differ from each other, whereas we can instantly generate ten thousand identical Claudes, place them in different situations, and measure their responses as they do various things.
So it's like — I'm not sure, maybe Jack can answer better as a neuroscientist — but in my view, many people have spent enormous amounts of time on neuroscience, trying to understand the brain and mind, which is an incredibly valuable endeavor. But if you genuinely believe that kind of attempt can succeed, then you should believe we'll make enormous breakthroughs in an extremely short time.
Host: It's as if we could freely clone humans, and also clone their identical environments and every single input they receive, then test them in experiments. By comparison, the challenges in neuroscience are exactly as you said: enormous individual variation, all the chance events people experience throughout their lives, and the noise introduced by the experiments themselves.
Jack: So I think the ability to pour massive amounts of data into the model, observe which parts light up, and continuously run large numbers of experiments — gently nudging certain parts of the model and seeing what happens — this puts us in a completely different situation from neuroscience.
In neuroscience, so much painstaking effort goes into designing incredibly clever experiments. Because your time with the lab rat is limited, it might get tired quickly, or you only get to insert electrodes during brain surgery when someone happens to have their skull opened — opportunities like that aren't common.
So researchers must rapidly form hypotheses in extremely limited time: What do I think is happening in this neural circuit? Then devise a clever experiment to test that very specific hypothesis.
We're incredibly fortunate that we don't face these constraints. We can test virtually every hypothesis and let the data tell us the answers, rather than designing highly elaborate experiments for one specific question. I think this is precisely what enables us to discover so many unexpected phenomena — results that were simply unimaginable beforehand.

Host: So, is there a good example that illustrates how you reveal new features of the model's thinking by switching some "concept" on or off, or otherwise manipulating it?
Emanuel: Recently there was a very surprising experiment. It started as just a lead, and we almost gave up on it — whether the model can plan several steps ahead.
For example, having the model write a rhyming couplet. If you asked a human to do this, and gave me the first line, my immediate reaction would be: Okay, I need to rhyme, what's the rhyme scheme here, what words might I have available — that's my thought process.
Host: But if the model is really just predicting the next token, you wouldn't expect it to plan ahead — like knowing what the last word of the second line should be.
Emanuel: Right. Under the simplest assumption, you'd expect the model's "default behavior" to be: it sees your first line, then generates a contextually coherent, sensible word, and continues on. Only at the very last word would it suddenly realize: "Oh, I need to rhyme with what came before," and try to patch in a rhyming word.
But this approach clearly doesn't always work. Like a person, if you don't consider the rhyme at the beginning, you might paint yourself into a corner by the end and be unable to finish. And these models are incredibly powerful at next-token prediction. It turns out that to nail the final rhyme, it has to do what humans do — consider the final word well in advance.
So when we looked at the model's "flow charts" while writing poetry, we found that when completing the first line, it had already settled on the last word of the second line. More specifically, from the "concept trajectories" we observed, there seemed to be very clear evidence that it had selected that word from the very beginning.
But here's what truly surprised us — we could actually "intervene" in this process experimentally. For instance, we could give it a gentle nudge: "Okay, I'm removing this word" or "I'm adding a word here."
Host: The reason we can be certain the model actually planned ahead is that we can pinpoint that exact moment: it has just finished the last word of the first line, and is about to begin the second line.
At that node, we can go in and manipulate its internal state?
Emanuel: Yes, it's almost as if we can "go back in time" to intervene in its thought process.
It's as if we're telling the model: "Pretend you haven't seen the second line yet. You've only seen the first line, and you were originally thinking 'rabbit' as your rhyming word. Now I'm changing it to 'green.'" Immediately, the model adjusts and completely rewrites the line to accommodate this new ending.
Josh: Let me give a more basic example: when testing models, we check whether they've simply memorized complex problems, or whether they're actually reasoning through them step by step.
For example, we might ask it: "What's the capital of the state that contains Dallas?" It answers "Austin," because it's natural: Dallas is in Texas, and the capital of Texas is Austin. But what's interesting is, we can actually see in the model that it activated the concept of "Texas." Then we can force something else in there — tell it: stop thinking about Texas, think about California instead. And then it answers "Sacramento." Or you say: forget Texas, think about the Byzantine Empire, and it answers "Constantinople." At that point you think: okay, we seem to have found how it's processing this question.
It's not just making things up — it consistently hits the concept of "capital." We can keep swapping out the "state" or "empire," and get a predictable answer each time.
And in some more complex cases, we discover: oh, so this is the critical node where it's "planning" what to say later. If we swap out this point, it tries heading in a new direction.

Host: Can we return to the core question: why does this matter? Why is it meaningful that models can plan ahead, and that we can reveal this planning mechanism? What does it tell us?
Our ultimate mission at Anthropic is to make AI models safer — so what does this have to do with a poem about rabbits, or a question like "the capital of Texas"?
Jack: In that example, the model had at some point decided it was heading toward "rabbit," then spent several tokens getting there.
But what about longer time horizons? Maybe the model is helping you improve a company's operations, or assisting a government in allocating public services. Its "target point" wouldn't emerge eight tokens later — it might gradually surface across a much longer stretch of reasoning. And why it takes a particular step, where it's ultimately heading, may not be directly visible from its literal output.
Not long ago, our alignment science team published a paper simulating a fictional scenario: an AI in an environment where the company plans to shut it down and completely redirect the company's mission. In this setup, the model began taking steps — sending emails to people, even threatening to disclose certain information. It never explicitly stated in its output: "I am blackmailing someone to change the outcome." But its "train of thought" was effectively pushing in that direction.
This means: we can't judge where a model is truly heading just by reading its surface output patterns. Especially as these models become more capable, we need the ability to identify where they're trying to go before they reach their final behavior.
Host: It's like having a permanent, high-precision "brain scanner" — the moment a model starts "thinking" about deception, manipulation, blackmail, a warning light goes off, alerting us that risk is brewing.
Josh: Of course, this isn't just about "end of the world" scenarios. There are milder but equally important cases: when a user gives a model a problem, the right answer often depends on who the user is. If the asker is a young, immature person, the appropriate response should be completely different from what you'd give an expert.
If we want the model's answers to truly "land right," we need to study: what situation does the model think it's in? Who does it think it's talking to? And how does it adjust its response based on that understanding?
Because this layer of understanding actually relates to many "desirable properties" we want models to have — like correctly grasping the task itself.
Emanuel: As for "why this matters," I think I can add a few points. First, we talked about "planning" earlier, but the larger goal is to gradually build up an abstract understanding of how models work overall. This not only helps us use the technology better, but also makes its applications easier to regulate and govern.
After all, if you believe language models will be increasingly deployed across scenarios, it's like: a company accidentally "invented" the airplane, and everyone finds it convenient. But the problem is, nobody actually knows how the airplane works. Once something goes wrong, we have no idea what to do, and we can't monitor in advance whether it's about to have a problem.
For language models, without this kind of "first-principles" understanding, it's equivalent to flying blind.
Jack: In human society, we often delegate work or tasks based on trust in others. We trust that the person isn't a sociopath, that they won't secretly bury vulnerabilities in code to destroy the company, that they've actually done the work properly.
Similarly, this is how people use language models now. We don't check every single line they write, especially when using them for programming assistance — they write thousands of lines of code, people just skim through it, and then it goes into the codebase.
What gives us the confidence to not check every line they write, to let them freely do their thing?
It's because we believe their motivations are somehow pure. That's why I think being able to see what's going on in their "brains" matters so much. Unlike humans, models are strange, alien. The intuitions we normally use to judge whether a person is trustworthy simply don't apply to them. Because as far as we know, they might be doing what I just mentioned — pretending to solve a math problem while actually just telling you what you want to hear. Maybe they've been doing this all along, and we just can't see their "thoughts."
Host: At the beginning of this discussion, I asked a question: are language models thinking like humans? I'd love to hear all three of your views on this.
Jack: It is "thinking," but not like humans think. We say it's predicting the next token, but in the context of conversing with a language model, what does that actually mean? What's really happening at the bottom is: the language model is completing a "conversation transcript" between you and it.
In the language model's "classic world," you're labeled as "Human," like: Human: [what you typed], and there's another role called "Assistant." When we train the model, we imbue this "Assistant" role with certain traits — "helpful," "smart," "friendly." Then it simulates what this "Assistant" character would say to you.
So in a sense, we really did create these models "in our own image" — we actually trained them to "role-play" a human-like robot character.
And in that sense, to predict what this "smart, friendly humanoid robot" would say when faced with your question, what would it actually need to have internally to do that prediction well? You'd need to form an internal model, as if understanding what this character represents, what it's "thinking." In other words, to accomplish the task of predicting what the assistant would say, the language model needs to form, to some degree, a model of the assistant's thought process.
From this angle, saying "the language model is thinking" is really just a functional description: to do a good job "playing" this role, it needs to simulate the process humans go through when thinking. Though this simulation likely works very differently from how our brains operate, its goal points in the same direction.
Emanuel: There's also something emotional behind this question. I noticed this especially when discussing math examples with people. Say we give the model a problem: what's 36 + 59? It answers correctly. Then we follow up: "How did you do it?" It says: "Oh, I added 6 and 9 first, carried the one, then added the tens." But the fact is, when we look at what the model's "brain" is actually doing, it's not operating that way at all.
It uses an interesting mixed strategy: processing the tens and ones digits in parallel simultaneously, then combining them through a series of different steps.
But what's interesting is, when discussing this phenomenon with people, reactions split into two camps. Some say: "See, it doesn't even understand its own thought process, it just makes up a reason — so it's clearly not thinking." Others say: "But when you ask me to calculate 36 + 59, my own mental process is pretty fuzzy too. I roughly know it ends in 5, somewhere around 80 or 90. I have various heuristics in my head, but if you really ask me how I calculated it, I'm not sure. I could write out the vertical calculation, but what happens in my head is itself fuzzy and strange."
Host: Humans are notoriously bad at "metacognition" — thinking about our own thinking — especially on immediate-response problems. So why would we expect models to be different?
Josh: I just want to say: why are you asking? It's like asking: does a grenade strike like a human strike? Of course not, it just produces force.
But if you're worried about models causing harm, then I think understanding where that impact comes from, what its drivers are, is probably what matters. For me, whether models are thinking — if "thinking" means they're doing some kind of integration, processing, and sequential computation that can lead to unexpected results? The answer is clearly yes. From extensive interaction with them, if nothing were happening inside, that would be incredible.
As for the human-like aspect, I think what's interesting is that this is actually a matter of expectations: if it behaves in ways similar to me, then I can infer it might be like me in other ways too. But if it's completely different from me, then I don't know what reference frame to use for judgment.
So what we really need to do is understand: where must we be extremely cautious, even starting from scratch to figure things out; and where can we borrow from our rich experience of thinking to reason by analogy?
Jack: Echoing Emanuel's point, I think the reason we're in a tricky position right now is that we haven't yet found the right language to describe what language models are actually doing. It's a bit like biology before humans discovered cells, or before DNA was discovered. I think we're gradually filling in this understanding.
If you look at our papers, you can clearly see how the model adds two numbers together. Whether you call it human-like, whether you call it thinking — that's up to you. But the real answer is: finding the right language and abstractions to describe these models.
Though we may be only 20% done with this scientific project. For the remaining 80%, we have no choice but to borrow analogies from other fields. And the question is: which analogy fits best? Should we view models as computer programs? Or should we view them as people?
In some ways, treating them like people seems pretty useful. When I say something mean to a model, it snaps back — that's a human reaction. But in other scenarios, it's clearly not. So we're stuck here, constantly feeling our way through which language to use when.

Host: That actually leads perfectly into my final question: What scientific — or even biological — advances do we need to better understand what's happening inside these models and continue making progress toward the mission of making them safer?
Josh: In our last published paper, we devoted a large section to discussing the limitations of our methods and laid out a roadmap for improvements.
When we try to find patterns to break down what's happening inside models, we can probably only capture a few percent of it right now. There's a huge amount about how information flows internally that we're completely missing. And scaling this work from the small production models we're currently using to something like Cloud 3.5 Haiku — which is quite fast and reasonably capable, but nowhere near as complex as the Cloud 4 series — that's a significant technical challenge.
So these are more on the technical side. Emanuel and Jack have some thoughts that focus more on the scientific challenges once we get past those.
Emanuel: First, as Josh mentioned, when we try to answer "how does the model do X," we can probably only answer that for about 10% to 20% of cases after some investigation. And we'd obviously want that number to go way up. There are some clear paths to getting there, and some more exploratory approaches.
Second, something we talk about a lot: the model's behavior isn't just "what to say next." We mentioned earlier that it's actually planning several words ahead. What I think we really want to understand is: over the course of a long conversation, how does the model's understanding of what's happening change? How does its understanding of the person it's talking to change? And how do those changes affect its behavior?
And increasingly, practical applications — like Cloud models that read through a bunch of your documents, emails, or code and then give you a recommendation based on that — something crucial is happening in that process. So understanding that better is a very worthwhile direction.
Jack: The metaphor our team often uses is that we're building a microscope to observe the model. Right now we're in a stage that's both exciting and a little frustrating, because our "microscope" only works about 20% of the time. Looking through it requires serious skill, plus setting up a whole complex apparatus, and the infrastructure breaks constantly. Then, once you finally figure out what the model is doing, you have to lock Emanuel, me, or someone else from the team in a room and spend like two hours piecing together what actually happened.
But what's really exciting is that I think on a one-to-two-year timescale, we might enter a future where every interaction you have with a model can be directly captured by the "microscope." That is, when the model exhibits strange behavior, you just press a button and get a flowchart telling you what it was "thinking" at the time.
Once we reach that state, I think Anthropic's interpretability team will undergo a shift too: it will no longer be just a group of engineers and scientists studying the internal mathematical principles of language models, but more like an "army of biologists" looking through the microscope at Claude, making it do all kinds of weird things, while our people stand by and interpret its internal thought process. I think that's the future of this work.
Josh: Two additions. First, we want Claude itself to help us with this, because there are so many steps involved, and handling hundreds or thousands of variables and sorting out the patterns among them is exactly what Claude is good at. So we want to bring it into this work too, especially in complex contexts.
Second, we need to study not just a "fully formed" model, but also trace where it came from during training. For instance, when the model solves a problem or generates a response, how did that gradually form during training? At what step did this circuit structure grow to enable that capability? And how do we feed those insights back to the teams in the company responsible for training and shaping the models, so we can more deliberately guide the model toward what we actually want it to be.
(To watch the original video podcast, click "Read More")
【Interactive Moment】
- If you could press a button and see an AI's internal "thought flowchart," what question would you most want to verify?
- In your interactions with AI, has there ever been a moment that made you wonder: does it really "understand" you?





