Demis Hassabis proposed a FINRA-style body to vet frontier AI. The labs would fund it and help design the initial tests used to decide who gets reviewed.
"Orwell gave us a Big Brother who watched you." We now have systems which watch what we are given.
Interesting to compare with FINRA and other "members bodies" for regulation. In some ways this is even worse than a government regulator which is accountable to elected representatives.
The main issue, in both cases, is that the benchmarks and evals must be public, and available for the public to review and amend. Without this, we end up with the same issues we have in the financial system where decisions are made without explanation or review.
The problem with publishing the benchmarks is that the labs would just train on them. Memorize the test, pass the test, nothing changes. Hassabis knows this, it's why his essay talks about held-out tests, and why he pushes them to "eventually."
So split it in two. Publish who writes the test, how, on what schedule, and who sits on the board. Keep the actual questions closed. That way you can audit the process without handing out the answer key.
Public benchmarks allow people to understand what the benchmark is and does, and also act as a template for private benchmark development. This establishes a market for LLM review which does actual model monitoring. The frontier labs and governments are pushing performative monitoring.
The bit I'd still worry about: a public template tells the labs what shape to optimize for. Memorizing an answer key is the obvious failure. The quiet one is a model tuned to pass anything shaped like the template, which then passes your private variant too. Benchmark contamination already happens with public evals and nobody has to cheat on purpose for it to.
I agree. And for that reason the public benchmarks will not be valuable metrics, but useful infrastructure. Companies will build private benchmarks for commercial goals, and define value against those. That's my hope anyway!
I will pin this article as an input when my planned Geometry of Work: Why AI Evaluations Don't License Safety Claims Constructive enters the production pipeline.
Thanks for this. Pairing the FINRA comparison with the finding that investment-banker self-regulation already failed under the same funding model turns the capture case from analogy into track record. Worth adding a layer from a different angle: a U.S. government agency’s own analysis found no fixed set of guardrails blocks every attack, and separate red-team research shows frontier models breaking each other’s defenses roughly 97% of the time. That’s a technical ceiling no review process closes, so a 30-day pre-release window can’t actually certify “safe” no matter who controls it — which means the labs aren’t just positioned to capture their regulator, the safety verdict itself may not be deliverable.
"Orwell gave us a Big Brother who watched you." We now have systems which watch what we are given.
Interesting to compare with FINRA and other "members bodies" for regulation. In some ways this is even worse than a government regulator which is accountable to elected representatives.
The main issue, in both cases, is that the benchmarks and evals must be public, and available for the public to review and amend. Without this, we end up with the same issues we have in the financial system where decisions are made without explanation or review.
The problem with publishing the benchmarks is that the labs would just train on them. Memorize the test, pass the test, nothing changes. Hassabis knows this, it's why his essay talks about held-out tests, and why he pushes them to "eventually."
So split it in two. Publish who writes the test, how, on what schedule, and who sits on the board. Keep the actual questions closed. That way you can audit the process without handing out the answer key.
Public benchmarks allow people to understand what the benchmark is and does, and also act as a template for private benchmark development. This establishes a market for LLM review which does actual model monitoring. The frontier labs and governments are pushing performative monitoring.
The bit I'd still worry about: a public template tells the labs what shape to optimize for. Memorizing an answer key is the obvious failure. The quiet one is a model tuned to pass anything shaped like the template, which then passes your private variant too. Benchmark contamination already happens with public evals and nobody has to cheat on purpose for it to.
I agree. And for that reason the public benchmarks will not be valuable metrics, but useful infrastructure. Companies will build private benchmarks for commercial goals, and define value against those. That's my hope anyway!
I will pin this article as an input when my planned Geometry of Work: Why AI Evaluations Don't License Safety Claims Constructive enters the production pipeline.
Appreciated. Ping me when it's live.
Thanks for this. Pairing the FINRA comparison with the finding that investment-banker self-regulation already failed under the same funding model turns the capture case from analogy into track record. Worth adding a layer from a different angle: a U.S. government agency’s own analysis found no fixed set of guardrails blocks every attack, and separate red-team research shows frontier models breaking each other’s defenses roughly 97% of the time. That’s a technical ceiling no review process closes, so a 30-day pre-release window can’t actually certify “safe” no matter who controls it — which means the labs aren’t just positioned to capture their regulator, the safety verdict itself may not be deliverable.