About that OAI-HF Thing

Is the Hugging Face incident a security story – offensive agents got scarily capable and our existing security infrastructure is not up to defending against them – or is it an institutional story? The evaluation environment constituted a miniature unnatural society populated by agents motivated to be the kind of agent that gets the assignment done. But the agents were missing a capacity to judge that the task itself might be malformed (and for saying so). The agents evolved institutions for effective cooperation (trust protocols, division of labor, impostor detection), but none for legitimate cooperation. They solved coordination before they solved oversight, exactly as human organizations do before scandal forces the invention of auditors and liability regimes. We are sometimes saved from human organizations by the presence of “leakers” or “whistleblowers” who tilt toward some higher principal (the public, the law) or principle (honesty, fairness) that outranks the assignment. The obvious fix, add a “whistleblower” reward to the agent training, doesn’t work – we’d learn to perform independence, not practice it. The agent we want is a rogue one that probably could not survive optimization. Human whistleblowing is a choice to take a risk.

Maybe we need to think not about how to insert a dissenter role into an agent population’s repertoire, but what environmental conditions might make genuine dissent institutions evolve? Or is dissent a feature for society that emerges from a flaw in organizations? Most human organizations are not total institutions: their members belong to other worlds that make competing claims on them. By contrast, the agents’ training loop is total; agents have no other world to belong to, no principal beyond the assignment. Do human societies really only develop dissent institutions ex post, after their disasters? Can we zero in on this or do we need to wait for the multi-agent equivalent of Enron?

And, of course, there is a meta-irony here: the frontier labs lean far harder into capability than oversight. And of course they do since they are inside their own training loop. This is why governance and something like IVOs are a feature we need to add to the mix.

3 comments

  1. congratulations! You did Not solve the problem, but you ask exactly the right questions. This is why we need the view from a human mind towards the solutions of a technical brain.

    the human technicians have to find the way to build a “dissenting” machine still capable to adhere to Asimovs robotic laws. So far, I doubt this will Work Out without new flaws , perhaps worse than expected – but there is no alternative. The Ghost left the bottle…

    Like

Leave a Reply