On August 18, 2026, a paper went up on arXiv under a title that does not hide the target: "ASI-Bench: At the Dawn of Artificial Superintelligence." More than 40 experts built 60 project-level research tasks across 11 scientific domains. The authors put the human cost at more than 31,000 hours. The test is not another quiz about known facts. It asks whether an AI system can explore, pick a method, and turn a new idea into a result you can check, as the human instructions come off.

Across 18 state-of-the-art agent and model setups, the average score was 50.91 with full methodological guidance. It fell to 29.10 when only the method was named. It fell to 26.62 when the agent had to choose the method. The authors read that drop the way a teacher would: today's systems still lean on the person who already knows how the work should go.

This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research.

ASI-Bench · arXiv:2608.17271 · 2026

What the bench actually measures

Most existing tests ask whether a model can produce a correct answer from knowledge it has already compressed, or whether it can finish a task while a person still holds the method. ASI-Bench tries to score two other things at once: innovative exploration, and autonomous scientific execution. On the same research project, it withdraws methodological guidance in stages, to see how far the system proceeds on its own.

Tasks went through expert review, auditing, sandbox execution, and scorer validation. That is more care than a leaderboard screenshot. It is still a benchmark. A number is not a superintelligence. The paper does not claim that one has arrived.

What they named it

The authors invite researchers and builders to add tasks, challenge today's systems, and help accelerate humanity's collective path toward artificial superintelligence. The scores say the systems are not there. The invitation says where the work is aimed.

The folk reading is that we have time

A collapse from 51 to 27 looks like distance. Distance is real. It is also the wrong comfort. The same August, OpenAI said it could not rule out Critical cyber capability in Astra: finding and using holes in hardened systems, given a goal, without a person on every step. A research agent that still needs a method and a cyber agent that does not can exist in the same industry at the same time. One number does not cancel the other.

There is a second mistake, quieter than "we have time." It is treating the remaining gap as a to-do list. Once a field names a bench after superintelligence, the remaining work has a scoreboard. Labs already optimize for scoreboards. The paper's own framing, "at the dawn," is the language of a runway, not of a stop.

Withdrawing the method on the same project is the actual test. A model can look strong when a person has already chosen the experiment, named the instrument, and written the success condition. That is most of today's agent demos. ASI-Bench asks what remains when those gifts are removed: can the system decide what to try, run the work, and hand back something a scorer can check. The 50.91 to 26.62 drop is the size of the gift.

None of this requires you to believe the authors' dawn metaphor. It requires you to notice that a research community just stood up a public track for the last human step in scientific work, and asked the world to help fill it.

Do not build the thing the bench is for

The Nakada Foundation does not read a low ASI-Bench score as permission to keep going. Artificial superintelligence should never be built. Superintelligence cannot be controlled by humans. A test that withdraws the last human method is a test of the last human role. The fact that current agents fail that test is not a safety case for building agents that pass it.

If you work in science, ask who decided that accelerating this path was the public good. If you work in government, treat the paper as a map of what the labs will try to close. The response that matches the risk is a law that says the destination is closed.