Technology

Goodfire's Inside-Out Monitors Watch AI Agents From Within

Martin HollowayPublished 14m ago3 min readBased on 3 sources
Reading level
Goodfire's Inside-Out Monitors Watch AI Agents From Within
Photo by Joshua Aragon on Unsplash

Goodfire launched monitors for AI agents on October 8, 2026 that inspect a model's internal signals as it works, rather than scoring output alone. The first release is available to customers of Baseten, which hosts and runs AI models for other companies. TechCrunch

The first implementation is built for the open model Kimi K3. Goodfire describes small detectors, called probes, that read internal signals at each step of a run. A separate AI model takes a closer look only when a probe flags something. The company calls that two-stage design inside-out monitoring.

Standard checks, including output-only monitors and LLM-as-judge systems where one AI grades another, must parse full histories after the fact. That means trajectories, tool calls made to outside software, and reasoning traces. Per-step reads from inside move the check earlier, close to the model's internal activity before behavior turns into text or tool use. Cost then scales with the number of escalations, not with total tokens reviewed.

Baseten customers can enable detection for three behavior classes: offensive hacking, chemical and biological weapons misuse, and reward hacking, where an agent games its scoring system instead of doing the intended task. The policy control is narrow. Operators can set the automated response to logging the event, routing it for human review, or refusing the request.

Goodfire's test figures on Kimi K3 are explicit. Monitoring about 1,500 sessions cost roughly $51 with probes plus selective review. The same workload cost $233 when a cheaper AI model checked every step, and about $10,000 when a top-tier model did the checking. The probes caught 94% of malicious hacking sessions. They sent 8.7% of harmless sessions for a second look. Running four probes at once added less than 2% to the time it takes the model to start responding, a budget known as time-to-first-token.

The work extends Goodfire's interpretability platform, Ember. Ember decodes neurons inside a model to provide programmable access to its internal workings, in the company's description. Goodfire announced a $50M Series A in April 2025 to build that platform. Goodfire The launch follows a September effort around open-weight safety, in which Base Labs started an open-weight AI safety partnership with Hugging Face and Goodfire to build evaluation and monitoring infrastructure for open-weight models. TechCrunch

The broader context here is deployment economics for agents. Multi-step agents multiply the tokens, tool calls and states a monitor must consider. Full review of every step works in evaluation but breaks in production on cost and latency. The Baseten integration is worth watching because inference providers already own routing, logging and policy enforcement, so embedded probes would give operators a control point without a parallel monitoring stack. In my view, the key tests are whether recall holds beyond Kimi K3, stands up to adaptive evasion, and keeps human review queues manageable at that 8.7% referral rate. We have seen this pattern before with spam and intrusion detection, where a cheap filter plus focused review became standard once it proved portable. If probes transfer cleanly to other open models, this approach could give agent deployments a practical safety layer.