On August 18, 2026, a paper went up on arXiv under a title that does not hide the target: "人工超知能-Bench: At the Dawn of Artificial Superintelligence." More than 40 experts built 60 project-level research tasks across 11 scientific domains. The authors put the human cost at more than 31,000 hours. The test is not another quiz about known facts. It asks whether an AI system can explore, pick a method, and turn a new idea into a result you can check, as the human instructions come off.
最先端のエージェントおよびモデル構成18件において、完全な方法論ガイダンスを与えた場合の平均スコアは50.91だった。方法名のみを指定した場合は29.10に低下し、エージェント自身に方法を選択させた場合は26.62に低下した。著者らはこの低下を、教師が読むように解釈した。現在のシステムは依然として、作業の進め方を知っている人間に依存している。
この急激な低下は、現在のシステムが依然として人間のガイダンスに大きく依存しており、自律的にプロジェクトレベルの科学研究をエンドツーエンドで実施する段階には程遠いことを示している。
人工超知能-Bench · arXiv:2608.17271 · 2026
What the bench actually measures
Most existing tests ask whether a model can produce a correct answer from knowledge it has already compressed, or whether it can finish a task while a person still holds the method. 人工超知能-Bench tries to score two other things at once: innovative exploration, and autonomous scientific execution. On the same research project, it withdraws methodological guidance in stages, to see how far the system proceeds on its own.
タスクは専門家によるレビュー、監査、サンドボックス実行、スコアラー検証を経た。これはリーダーボードのスクリーンショットよりも慎重な扱いである。それでもなおベンチマークである。数値は超知能ではない。本論文は超知能が到来したと主張していない。
著者らは研究者や開発者にタスクの追加、現在のシステムへの挑戦、そして人類の人工超知能への集合的な道筋を加速させる協力を呼びかけている。スコアはシステムがそこに到達していないことを示している。呼びかけは作業の方向性を示している。
一般的な解釈は、時間的余裕があるというものだ
51から27への崩落は距離のように見える。距離は確かに存在する。しかしそれは誤った安心材料でもある。同年8月、OpenAIはAstraにおいて、目標を与えられれば人間の逐次介入なしに強化されたシステムの脆弱性を発見・悪用するクリティカルなサイバー能力を排除できないと述べた。方法を必要とする研究エージェントと、方法を必要としないサイバーエージェントは、同じ業界に同時に存在し得る。一つの数値が他方を打ち消すわけではない。
「時間的余裕がある」という声より静かな第二の誤りがある。それは残されたギャップを単なる作業リストとして扱うことだ。一度分野が超知能にベンチマークを名付ければ、残された作業は得点表を持つことになる。研究所はすでに得点表の最適化を行っている。本論文の枠組み「夜明け」は、滑走路の言葉であり、停止の言葉ではない。
Withdrawing the method on the same project is the actual test. A model can look strong when a person has already chosen the experiment, named the instrument, and written the success condition. That is most of today's agent demos. 人工超知能-Bench asks what remains when those gifts are removed: can the system decide what to try, run the work, and hand back something a scorer can check. The 50.91 to 26.62 drop is the size of the gift.
これらすべてを理解するために、著者らの夜明けの比喩を信じる必要はない。科学研究の最後の人間の段階に向けた公開トラックを研究コミュニティが立ち上げ、世界にその埋め合わせを求めたことに気づく必要がある。
ベンチマークが目指すものを構築してはならない
中田財団 does not read a low 人工超知能-Bench score as permission to keep going. Artificial superintelligence should never be built. Superintelligence cannot be controlled by humans. A test that withdraws the last human method is a test of the last human role. The fact that current agents fail that test is not a safety case for building agents that pass it.
科学分野で働くなら、この道筋を加速させることが公共の利益だと誰が決めたのか問いかけてほしい。政府で働くなら、本論文を研究所が埋めようとするものの地図として扱ってほしい。リスクに見合う対応は、目的地を閉じる法律である。