Static

Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

First reported by TechCrunch ·

The signal ●○○○ Compiled by AI from TechCrunch, the single source so far
Why you might care

AI monitoring costs fall by up to 98%, making it practical to guard open models against misuse.

What happened

Goodfire, a startup specializing in AI interpretability, has launched a new method for monitoring AI agents, dubbed ‘inside-out’ monitors. This system observes an AI model's internal calculations during its operation, rather than relying on a secondary AI to review the output. This approach aims to be significantly cheaper and more efficient than existing methods, which involve a separate AI rereading all processed text. Goodfire's probes tap into computations the AI is already performing, thus reusing existing processing power. The company claims this reduces monitoring costs dramatically, with a test on the Kimi K3 model costing around $51 for 1,500 sessions, compared to thousands for traditional methods. The system can detect risks like malicious hacking and misuse of sensitive information, offering configurable responses from logging events to halting requests. This technology is primarily targeted at open AI models, which often lack the built-in safeguards of proprietary systems.

What it means

Goodfire's 'inside-out' monitoring approach represents a significant shift in AI safety by making comprehensive internal process observation economically viable. By leveraging existing computational steps within the AI model, the cost and latency associated with real-time safety checks are dramatically reduced. This innovation could democratize advanced AI security, bringing robust safeguards to open-source models that were previously too expensive to monitor effectively.

This development is particularly impactful for the open AI ecosystem, where rapid iteration and deployment have outpaced robust security measures. As AI agents become more sophisticated and their potential for misuse grows, Goodfire's cost-effective solution addresses a critical market need for both developers and inference providers. The ability to detect and prevent malicious activities at a fraction of the cost may encourage wider adoption of AI while mitigating associated risks.

AI-written summary. May contain errors.