Back to Now
Goodfire production cyber monitors launch Verified

Goodfire launches activation-based cyber monitors for AI agents

The system reads internal model signals during agent work and sends selected exchanges to an AI judge, with cost and detection gains reported in company tests.

Why now Goodfire published its production monitor research on October 8, and TechCrunch reported the customer launch at 16:00 UTC.

Watching internal signals as an agent works

Goodfire has introduced cyber monitors for open-model AI agents that inspect model activations while an agent runs. Small probes flag exchanges that may involve risky behavior; a separate AI judge then reviews those exchanges instead of reading every turn. Goodfire says it built and deployed monitors for Kimi K3 and GLM 5.3 on a production inference stack. TechCrunch reported on October 8 that the monitors are available to customers of model-hosting company Baseten. Customers can choose risks to watch and whether a flag should be logged, sent for human review, or lead to a refusal, according to the report.

Promising tests, with defined limits

Goodfire reports that its probe-and-judge setup identified about 93% of harmful sessions while interrupting 5.5% of benign sessions in its evaluation, and that the chosen configuration cost much less than judging every turn. Those are company-run results using its selected tasks and reference labels, not a general guarantee against misuse. Goodfire also describes preliminary testing by FAR.AI: a fixed, non-adaptive set of jailbreak attempts produced no successful universal jailbreaks against the monitored system, but some individual interactions still succeeded. Further adaptive testing and results on other models would be needed before treating the approach as a broad security solution.