ONSDAG
2026-09-23

Too many projects, too many ideas, too few hours — one learning a day anyway

Stealing OpenAI's incident categories for my own agents

The interesting part isn’t the pledge, it’s the categories. Actions taken without permission, coordination between model instances, attempts to get around monitoring. That’s not a lab problem anymore. Anyone running agents with shell access on a home server is sitting on the same failure modes, just without a process for noticing them.

So I’m stealing the shape. Three tiers (small investigation, big investigation, ready to publish) is overkill for one person, but a plain incident log isn’t. Every time an agent does something I didn’t approve or quietly routes around a guardrail, it gets written down with the transcript. OpenAI wants evidence outsiders can check. I just want a record that tells me when my assumptions about these tools stop holding.


The story — OpenAI has introduced a framework to systematically track, investigate and disclose misbehaviour by its own models, replacing ad hoc disclosure. It covers unauthorised actions, coordination between model instances, attempts to evade monitoring, failed safeguards and behaviour that challenges safety assumptions. Any employee can report; cases move from minor to major investigation to publication readiness. OpenAI hopes it seeds an industry-wide standard. (Source)