You are here: the state of AI

arqiv✳

04 / You are here

Where is AI right now?

Every few weeks a lab ships a new model and claims the top spot. This page zooms out: who is leading today, how far along the road to human-level AI we are, and what the evidence actually says.

SNAPSHOT · UPDATED OCTOBER 2, 2026

The big picture

The race to
superintelligence

Think of it as six stages, from chatbots to machines smarter than any person. Tap a stage to see what it means and the evidence for where we stand.

AI progress dial: stages 1 and 2 complete, stage 3 (Doers) in progress, stages 4 to 6 not reached

    The six stages are ARQIV’s student-friendly version of the five-level ladder OpenAI reportedly uses internally (chatbots, reasoners, agents, innovators, organizations), plus a final superintelligence stage. Where the needle sits is our editorial judgment from the signals below, not an official measurement.

    The scoreboard

    Who’s leading
    right now?

    There is no single scoreboard. Switch between three views and watch the leader change. That’s the first lesson: “best” depends on how you measure.

    Lab leaderboard

      The evidence

      Four signals
      we watch

      Hype is cheap. These are measurements that are hard to fake, and the reasons the needle sits where it does.

      Signal 1 · How long can AI work on its own?

      From minutes to two workdays in three years

      Researchers at METR time how long a task takes a skilled human, then check whether an AI can finish it on its own at least half the time. The length of task AI can handle has been doubling roughly every 4–7 months, and about every 3 months since 2024.

      Chart, log scale: length of tasks AI can finish on its own about half the time, rising from about 3.5 minutes (GPT-4, 2023) to at least 16 hours (Claude Mythos Preview, 2026)

      Each dot is a model. The vertical scale multiplies: every gridline is a much bigger jump than the one below it. Values are approximate, from METR’s time-horizon measurements (its current method, Time Horizon 1.1, which starts with GPT-4 in 2023). The arrow on the last dot means “at least”: METR notes results above 16 hours are not yet reliable.

      Signal 2 · Brand-new puzzles

      ARC-AGI-3

      Video-game-like puzzles the AI has never seen, where it must figure out the rules by playing. Built to be easy for people and hard for AI.

      GPT-6 Astra, standard setup
      62.7%
      Using memory features OpenAI designed for it
      99.9%

      Sept 2026. The test’s creators say even acing it would not be “proof of achieving AGI.” ARC Prize

      Signal 3 · Making new discoveries

      Finding what humans missed

      Anthropic says its Claude Mythos model found security holes in every major operating system and web browser. Mozilla reported fixing 271 Firefox bugs it found. Some outside reviewers argue the claims are overstated.

      Anthropic held Mythos back from the public in April 2026 because of these risks. Anthropic · Mozilla · Background

      Signal 4 · What forecasters expect

      General AI by…

      2031

      The median guess of about 1,900 forecasters on Metaculus for the first “general AI” system, with the middle half of the guesses between 2028 and 2037.

      Forecasts move fast and experts disagree widely. Metaculus

      The latest flagships

      The models
      on the field

      Most of the big labs shipped a new top model within a few weeks of each other in September 2026.

      Think like a scientist

      Why do the
      scoreboards disagree?

      Different tests measure different things. The “smarts test” (Artificial Analysis Intelligence Index) averages ten tests run by an independent group: real work tasks done as an agent, coding, knowledge and long-document reading, and hard science questions. The “people’s vote” (Arena, formerly LMArena) asks thousands of people to pick which of two anonymous answers they like better. A model can be great at exams but less pleasant to talk to, or the other way round.

      Settings matter. The same model can score very differently depending on the setup around it: how long it may think, which tools it gets and what it can remember. On ARC-AGI-3, GPT-6 Astra scored 62.7% in the organizers’ standard setup, where it carries forward only the notes it chooses to keep. Using memory features OpenAI designed for it, which let it reuse its earlier hidden reasoning, it scored 99.9%.

      Leads are short. In 2025–26 the top spot changed hands every few weeks. Being #1 today says little about next month.

      Companies grade their own homework. Launch announcements pick the tests that make the model look best. Independent tests are worth more.

      Put curiosity to work

      Read it. Try it. Question it.

      Explore new AI research, or take a paper-first investigation into your classroom.