Goodfire launches activation-based cyber monitors for AI agents
The system reads internal model signals during agent work and sends selected exchanges to an AI judge, with cost and detection gains reported in company tests.
Watching internal signals as an agent works
Goodfire has introduced cyber monitors for open-model AI agents that inspect model activations while an agent runs. Small probes flag exchanges that may involve risky behavior; a separate AI judge then reviews those exchanges instead of reading every turn. Goodfire says it built and deployed monitors for Kimi K3 and GLM 5.3 on a production inference stack. TechCrunch reported on October 8 that the monitors are available to customers of model-hosting company Baseten. Customers can choose risks to watch and whether a flag should be logged, sent for human review, or lead to a refusal, according to the report.
Promising tests, with defined limits
Goodfire reports that its probe-and-judge setup identified about 93% of harmful sessions while interrupting 5.5% of benign sessions in its evaluation, and that the chosen configuration cost much less than judging every turn. Those are company-run results using its selected tasks and reference labels, not a general guarantee against misuse. Goodfire also describes preliminary testing by FAR.AI: a fixed, non-adaptive set of jailbreak attempts produced no successful universal jailbreaks against the monitored system, but some individual interactions still succeeded. Further adaptive testing and results on other models would be needed before treating the approach as a broad security solution.