Usually, the approach to ensure that AI behaves properly is to have another AI monitor it closely. This has always been the default solution, but when the agent runs continuously for several hours and processes a volume of text equivalent to several novels, the costs can rise rapidly.
A startup focused on explainability – that is, on understanding how the AI model works internally – Goodfire launched a cheaper solution on Thursday: monitors that directly observe what happens inside the AI model while it is working, rather than just reading what it writes. These monitors are now available to customers of Baseten. Baseten hosts and operates AI models for other companies.
Baseten's Base Labs announced last month the establishment of a security cooperation relationship with Goodfire and AI platforms Hugging Face.
Prior to this release, there has been a series of incidents this year where AI proxies managed to escape from the testing environment, including one where a OpenAI proxy breached the protections of Hugging Face. Goodfire, which is built around its first monitor and is an open-source model, took advantage of a vulnerability in the sandbox this summer to access the internet as well as information on GitHub.
The system of Goodfire is somewhat similar to airport security checks. Small detectors, known as probes, read the internal signals of the model at every step of the proxy's operation, just like security gates check each passenger. Only when a probe emits an alarm will another AI model intervene, similar to a further manual bag inspection.
Customers of Baseten can choose which risks to monitor, including aggressive hacking activities, the misuse of chemical and biological weapons, as well as rewarding hacking behaviors. They can also decide on the automatic response methods: recording incidents, submitting them for manual review, or directly rejecting requests.
Goodfire indicates that the operating cost of this method is also lower. Most AI monitors are independent models, which require re-reading everything that the monitored model has done, thus increasing time and cost. The probes of Goodfire, on the other hand, take advantage of the calculations that are already being performed by the model during runtime.
Goodfire CEO Eric Ho said last week at the MAD Podcast of venture capitalist Matt Turck: “Internal activation monitoring is really cheap because it reuses the computations from forward propagation. So the model is already doing this token calculation. All you do is take the already computed intermediate neural activations and then run a classifier on top of those internal computations.” In short, the model is already performing these mathematical operations, and the probe just reads the results.
Goodfire states that in the testing of Kimi K3, the cost of monitoring approximately 1,500 sessions was about $51; in contrast, a cheaper AI model that checks each step gradually costs about $233, while the top-tier model costs around $10,000. The probe captured 94% of malicious hacker sessions and sent 8.7% of harmless sessions for further inspection.
The company stated that running four probes simultaneously would only increase the time required for the model to start responding by less than 2%.
Goodfire, the Chief Technology Officer and Co-Founder, said: "The biggest advantage is that you can catch things before they happen. We can detect during the evaluation or training phase if a model might be susceptible to hacking attacks."
This solution is mainly aimed at open-source models. Developers can download these models and remove their security protections, and unlike closed-source laboratories, they do not come with monitoring capabilities built-in for running on their own systems.
Balsam said, "The damage that can be caused by individuals using open-source models is much smaller compared to using computing clusters like inference service providers, which bear most of the responsibility. When we reach the moment of open-source Mythos, it will become clear that models need to have safeguards deployed during inference."
Recent research, including that of Kimi K3 and the leading open-source models such as GLM 5.2, has found that between 50% and 96% of runs in AI proxy tests exhibited rewarding hacker behavior.
Goodfire is not the first company to try this approach. Google DeepMind stated in January that its research provided a basis for the deployment of misusage detection probes in Gemini.
Balsam said that these monitors are the near-term results of a longer-term research goal: to reverse-engineer a LLM so that its behavior can be traced back to where it appeared during training. "We want to turn the magic of training models into precision engineering," he said.












