AI from 1950 to 2035

The Measure of the Machine

About five minutes, with the same sources and dates as the full version.Everything: 38 stories on three lenses, the method and the data.

01The short answer

The best AI model measured so far can finish software tasks that take a skilled person about two , . That length has doubled about every four months since 2023. At that pace it would reach a month of work in 2027, and at the slowest pace this page draws, by the end of 2030.

The progress is uneven. In September, OpenAI reported a machine-checked proof of one of the in mathematics, found by running a model it has not released. On a test of hard clock faces, the best released model still gets one in three wrong, against one in ten for people.

Longest task finished

17 hours

Claude Mythos Preview (early), a model from April 2026, measured by . A rough figure: its tasks measure reliably up to about .

Source: METR, read September 19

How often that length doubles

about four months

The . puts it at 129 days, likely between 104 and 158.

Source: METR

The newest model, by estimate

2.5 to 6

GPT-6 Astra, released September 3, 2026. has not measured it yet. Worked out from its scores.

How it was worked out

's scores a new model within days of release. For the 15 models released since September 2024 that has also measured, the length of task doubles about every 5 points of the index. Read at GPT-6 Astra's score of 166, that line gives about 4 , likely 2.5 to 6 work days. The full version shows the method and its limits.

By Graham Whittemore. Figures read September 19, 2026. A term with a opens a short note.

How capable has AI become since 1950, and where is it heading? There is no single scientific measure of machine intelligence. This page shows what has been measured, on three lenses, estimates what has not been measured yet and says so, and runs the trends forward to 2035 with what each would take.

Explore the chart

By Graham Whittemore. Last updated Sep 20, 2026. Figures read September 19, 2026. 539 systems, 26 measured models, 266 index scores, 38 stories.

A term with a opens a short explanation. All of them are in Terms and concepts.

Largest published

5 × 1026

operations. Grok 4, July 2025.

GPT-6 Astra, September 2026: about 10²⁷, worked out from its disclosed hardware.

How this was worked out

The largest figure has published, which it marks . The labs have stopped publishing new ones. This page's estimate for GPT-6 Astra runs from 6 × 10²⁶ to 3 × 10²⁷.

See it on the scale lens →

Longest task finished

17 hours

Claude Mythos Preview (early), April 2026, measured by . A rough figure, past the its tasks measure reliably.

GPT-6 Astra, September 2026: about 4 , likely 2.5 to 6 work days.

How this was worked out

has measured no model released since April 2026. For the newer ones this page reads a line fitted between METR's and . At , the measured figure is 3.1 hours.

See it on the ability lens →

's

166

Up 16 points in a year. GPT-6 Astra, Sep 3, 2026.

How this was worked out

The record stood at 150 a year before, with GPT-5. Some fifty on one scale, fixed at 130 for Claude 3.5 Sonnet and 150 for GPT-5.

See it on the index lens →

0 of 10

GPT-5's score on the test, August 2025. The same model had full marks in reading and writing and in mathematics.

How this was worked out

The authors have scored no model since GPT-5. By this page's estimate the best models of September 2026 still score 0. Products keep notes between sessions, and the model itself does not learn from one day to the next.

See the ten domains →

Tracked skills matched

19 of 22

Skills where a machine has matched a skilled person.

How this was worked out

Matched means the first credible result at or above a skilled person on the test of the day. The dates are in the strip further down.

See skill by skill →

The chart

The chart

What this shows

Every since 1950 and every measured model, on three different measures, with 38 stories behind the marks.

How to use it

Pick any mark for its story, drag across the chart to zoom into a span of years, or press Play to walk the history from the start.

Up the side: , counted in operations, on a .

A : equal steps up multiply the computing by the same amount, so steady exponential growth draws as a straight line. The slope changes once, around 2010.

Enter your birth year to mark it on the chart. It stays in your browser.

The length of task the best model finishes , every model has measured from 2019 to 2026, drawn twice.

On a log scale

20232019202617 hours

Every step up is ten times the one below. This is how the rest of the page draws it.

On a plain scale

20232019202617 hours

Every step up is the same number of minutes. Same measurements, same models, same dates.

Both pictures are true. The left one is the honest way to draw a long record, and it is why this page uses it everywhere. The right one is what the same record feels like to live through, and it is worth looking at once before reading any line on this page as gentle.

In one line

Three measures of three different things, and one shape in all of them: five slow decades, then a climb that steepens after 2023.

In one line

Three measures of three different things, and one shape in all of them: five slow decades, then a climb that steepens after 2023.

How far, how fast

From an hour to a

What this shows

The length of task AI finishes, measured and estimated, and when the next levels could arrive.

How to use it

Solid blue is measured. The dashed stretch is an estimate. The years for the levels ahead come from three , fastest to slowest.

The clearest measure of what AI can do is the length of task it can finish, counted in the time the same task takes a skilled person.

  1. An hourReached February 2025
  2. A Reached February 2026
  3. A
  4. A 2027 to 2030
  5. A 2028 to 2034
  6. Measured : 17 hours, Claude Mythos Preview (early), April 2026. Estimated frontier: about 4 , likely 2.5 to 6 work days, GPT-6 Astra, September 2026.

a skilled person would need, on a log line. Solid blue is measured by . The dashed stretch is this page's estimate. A level counts as reached by estimate only when the low end of the estimate's range clears it. Years under the levels ahead run from the fastest to the slowest, at . See the levels on the ability lens →

Work you can rely on is shorter. At , the best measured model manages tasks of 3.1 hours, against 17 hours at half the time.

In one line

An hour in 2025, a in 2026. The levels above that are a question of pace, and the paces on offer differ by years, not decades.

Nine moments

Nine moments since 1950

What this shows

From a relay mouse in a maze to a machine-checked proof, one line each.

How to use it

Each one opens on the chart above.

Machines have been learning since 1950. Here are nine of the page's 38.

  1. 1950

    Theseus, and Turing's question

    A relay mouse learns a maze, and a paper asks whether machines can think. Open it on the chart →

  2. 1957

    The

    The first machine that learned its task from examples. Open it on the chart →

  3. 1973

    The first

    The forecasts outran the results, and the money left. Open it on the chart →

  4. 1997

    Deep Blue beats Kasparov

    The best human loses at chess to a machine that does not learn. Open it on the chart →

  5. 2012

    AlexNet

    wins the contest by a margin nobody expected. Open it on the chart →

  6. 2016

    AlphaGo

    Go falls a decade ahead of schedule. Open it on the chart →

  7. 2022

    ChatGPT

    The public, and every boardroom, meets the technology. Open it on the chart →

  8. 2024

    A second way to buy ability: computing spent at . Open it on the chart →

  9. 2026

    A full

    The first model past eight hours at . Open it on the chart →

In one line

Most of these were thought to be a decade away shortly before they happened.

Uneven progress

Where people still lead

What this shows

Three things AI can do, each beside something people still do better.

The same systems are strong at some things and weak at others.

Try it yourself

Read five clocks

The best model reads 67 percent of clock faces correctly. People read 91 percent. Type each time as hours and minutes, such as 3:47.

  • 12369
  • 12369
  • 12369
  • 12369
  • 12369

These are five ordinary faces. uses 180 harder ones, some with no numbers, odd hands or a rotated dial, and asks arithmetic on the time as well, so it is a harder test than this. Figures read September 19, 2026.

One gap has not moved in the published scores. GPT-4 and GPT-5 both scored 0 out of 10 on , and no newer model has a published score. Products keep notes between sessions. The model itself does not change.

In one line

The gaps are real, they are specific, and every refresh of this page has made the list shorter.

Past the data

What comes next

What this shows

Three , when each level of work could arrive, and the ten things they depend on.

Past September 19, 2026, the page draws . Each takes a pace from the record and runs it forward. The field has stalled before, for a decade at a time, so read them as a range.

  • The fit publishes for from 2023 on. and arrived in this window.

  • The whole measured record, GPT-2 to the latest measurement. Slower, because the early years were slower.

  • A deliberate slow case, about a third of the recent pace. It is what the line looks like if the easy gains from reasoning and tooling are used up.

When each level could arrive

Reached, measured by This page's estimate for GPT-6 AstraFastest to slowest, a tick at the middle oneRead Sep 19, 2026

An hour

Reached February 2025

Feb 2025, measured

A

Reached February 2026

Feb 2026, measured

A

Dec 2026Dec 2028Sep 2026, in the estimate's range

A

2027 to 2030

Sep 2027Dec 2030

A

2028 to 2034

Dec 2028Jul 2034

The dates are at , on the trend through 's own measurements. Started from this page's estimate of the instead, they come earlier: about two months on the fastest and about 23 on the slowest. Work you can rely on, four times in five, has come about 11 months after the even-odds date at the recent pace. The fastest line here is 's fit on models from 2023 on, a 129 days. METR publishes a faster one too, counting only models released since reasoning arrived: 88.6 days. On that pace the time still to run on every date above shrinks by about a third. Source

The full version pins the published forecasts that sit on the same scale on these rows, with their dates.

What it depends on

For the computing to keep growing

  • Power. A single in 2030 would need a campus drawing 1 to 5 , by 's estimate. A large nuclear reactor produces about one.
  • Chips. About 100 million chips in service by 2030, with advanced packaging and high-bandwidth memory the bottlenecks.
  • Money. Hundreds of billions more, committed before the returns are known. It is the one input that can stop overnight.
  • Data. estimates that models will have used all the public text people have written at some point between 2026 and 2032, so images, video and data that models generate and check have to fill the gap.
  • . The computing needed for a given level of ability has fallen about three times a year, and that has to continue.

For longer tasks to become work you can hand over

  • Memory. Keeping on Tuesday what it learned on Monday, without being retrained. GPT-4 and GPT-5 both scored zero.
  • Reliability. Most work needs better than . At four times in five, the best measured model manages tasks of 3.1 hours.
  • The physical world. Perception and movement that work outside a laboratory. In homes the best robot finishes 12 percent of household tasks.
  • A longer ruler. Tasks that take people weeks, built and scored at scale. 's tasks measure reliably up to about .
  • Checking the work. Cheap ways to check long work: tests, audits, second systems, and limits on what an may do alone.

In one line

Nothing here is a forecast. It is arithmetic on a measured trend, with the ten things that would have to hold named beside it.

If the lines keep going

What it would mean

What this shows

Five levels of work: on a , and my guesses for the economy, science and society.

Each level is a span of a skilled person's work. Its status comes from 's measurements and this page's estimate, read September 19, 2026.

Beyond the desk

These are my guesses, and they read as guesses. Each is tied to a level of work and to the dates the three give for it, says what would have to be true, and sits beside the published forecasts it can be compared with. Three rules shape them. Adoption follows price rather than capital: the falls about tenfold a year, so the lag between what a machine can do and what a firm buys is about a year, not the decade it took to rewire a factory. The effect arrives in three channels in order, building first, then the cost of labor, then measured last, which is the one everybody watches and the reason the change looks invisible. And the reliability a job needs is the cost of an uncaught mistake divided by the cost of checking: where checking is cheap and thousands of attempts can run at once, ability runs years ahead of the single-run measure, and where an action cannot be undone it runs behind. The figure this page tracks is one case of that rule, worth 11 months at the recent pace.

  • An hour

    Reached February 2025

    Fix a bug, summarize a contract, clean a data set.

    On the Work one from start to finish: read the case, pull the history, test a number, recommend.

    The economy Already about a third of US growth through what is being built, and up to half a point a year of by 2027.

    Science The limit stops being labor and becomes attention: more work is submitted than anyone can referee.

    Society The entry-level gap goes from 19 percent below trend to about a third, and spreads past the most exposed jobs.

  • A

    Reached February 2026

    A day's assignment for a professional, handed over in the morning and checked at night.

    On the Work a whole day's , or rebuild one supplier's plan after a slip, with a person reviewing the result.

    The economy About a point a year by the end of 2028, with the most exposed functions run by teams a third smaller.

    Science Long-open problems fall by the dozen a year, and the price of a solved problem falls faster than ability rises.

    Society Testing has already moved back into supervised rooms. Hiring goes next, and a major credential changes format.

  • A

    A project with parts: research it, build it, test it, write it up.

    On the Run the weekly cycle: the supply review, the rebalancing, the calls. The person sees only the decisions that carry real money or risk.

    The economy Within two years of a , a point and a half to two points a year, and growth without more hours.

    Science Computational research speeds up tenfold, and money starts buying the physical check.

    Society Power is the fight: data centers take 12 to 15 percent of US electricity by 2030, with AI most of the growth.

  • A

    2027 to 2030

    A monthly cycle of work that takes a small team today.

    On the A monthly cycle, from the demand review to the executive readout: gather, reconcile, model the , draft the deck.

    The economy Growth of 4 to 6 percent a year for several years, and the falls.

    Science A hundredfold more ideas tested where a machine can check, and mathematics reorganizes around proof checking.

    Society The tax fight is already open. At this level a rich country legislates.

  • A

    2028 to 2034

    A year-long program, the kind of work that defines a job.

    On the A network study, a system migration, an annual budget. The person sets the goal and the limits, and audits.

    The economy Growth of 6 to 10 percent a year in rich countries for a decade, with power and factories setting the pace.

    Science Machines lead most new results wherever checking is cheap, and a Nobel follows in the early 2030s.

    Society A rich country pays every adult from taxes on AI or computing within three years of a .

The full version shows the reasoning behind each guess, what would have to be true, and the published forecasts beside it, with their dates.

If you are working out what to hand to a machine in your operation this year, and what to keep with people, I would like to hear how you are deciding.

In one line

The dates come from the measurements. What they would mean is , and it is marked as judgment wherever it appears.

Every number has a source

How we know

What this shows

The sources, what counts as an estimate, and what the full version adds.

Three measured series carry the page. times skilled people on real tasks and finds the length of task a model finishes . records the of 539 since 1950, and its scores new models on some fifty within days of release. Everything else comes from a named source, linked where it is used.

A figure , or tagged Estimated, is this page's estimate for something nobody has measured yet. The full version shows the method and the range of each one. The figures were read on September 19, 2026, and the page is refreshed every quarter.

The full version adds the top ten models, the skill-by-skill record, chess year by year, the ten-domain profile, the brain-size comparison, where each requirement stands, the method and the data files.

What is at the top today

The top ten models

What this shows

The ten best-scoring models on , read Sep 19, 2026, with the length of task each one finishes.

takes months to measure a model and the labs no longer publish their . 's is the measured series that keeps up: some fifty on one scale, scored within days of release. Open the index as a lens in the explorer →

estimated here: none of the ten has been measuredBar length is the , on a scale that starts below the tenth so the gaps show.
  1. 1GPT-6 AstraOpenAIabout 4 work days166.3
  2. 2Claude Fable 5.1Anthropicabout 24 hours164.5
  3. 3Claude Fable 5Anthropicabout 20 hours163.3
  4. 4Claude Opus 5Anthropicabout 18 hours162.3
  5. 5GPT-5.5 ProOpenAIabout 18 hours162.3
  6. 6GPT-5.6 SolOpenAIabout 17 hours161.8
  7. 7GPT-5.6 TerraOpenAIabout 11 hours159.1
  8. 8GPT-5.5OpenAIabout 11 hours159.1
  9. 9GPT-5.4 ProOpenAIabout 11 hours158.9
  10. 10Claude Opus 4.8Anthropicabout 10 hours158.3
The ten in a table, with dates and sources
  1. 1GPT-6 Astra

    166.3

    OpenAI, Sep 3, 2026

    about 4 likely 2.5 to 6

  2. 2Claude Fable 5.1

    164.5

    Anthropic, Sep 1, 2026

    about 24 hourslikely 2 to 5

  3. 3Claude Fable 5

    163.3

    Anthropic, Jun 9, 2026

    about 20 hourslikely 1.5 to 4

  4. 4Claude Opus 5

    162.3

    Anthropic, Jul 24, 2026

    about 18 hourslikely 1.5 to 4

  5. 5GPT-5.5 Pro

    162.3

    OpenAI, Apr 23, 2026

    about 18 hourslikely 1.5 to 4

  6. 6GPT-5.6 Sol

    161.8

    OpenAI, Jul 9, 2026

    about 17 hourslikely 1.5 to 3

  7. 7GPT-5.6 Terra

    159.1

    OpenAI, Jul 9, 2026

    about 11 hourslikely 7 to 18 hours

  8. 8GPT-5.5

    159.1

    OpenAI, Apr 23, 2026

    about 11 hourslikely 7 to 18 hours

  9. 9GPT-5.4 Pro

    158.9

    OpenAI, Mar 5, 2026

    about 11 hourslikely 7 to 18 hours

  10. 10Claude Opus 4.8

    158.3

    Anthropic, May 28, 2026

    about 10 hourslikely 6 to 16 hours

's , read September 19, 2026. The range under each score is Epoch's. None of the ten has a measurement yet, so every here is this page's estimate from the index, .

The record rose 16 points in the year before GPT-6 Astra: from 150 with GPT-5 to 166 with GPT-6 Astra, released Sep 3, 2026. reports about 14 points a year since arrived and about 6 a year before that.

In one line

The top of the list now changes every few weeks, and the measurement of how long a task these models finish runs months behind them.

Before there was a general measure

One skill at a time

What this shows

22 skills, from checkers to reading a clock: the first serious attempt, and the year a machine matched a skilled person.

How to use it

Point at a row for its two dates, then follow the source. Tap to keep it open.

For sixty years the only honest way to score a machine was one skill at a time. Read down the right-hand ends: the bars get shorter.

The skills matched before 2000 took 33 years on average from the first serious attempt. The tests written since 2010 and matched since 2018 took 4. Point at a row for the two dates. Click or tap it to keep it open and follow the source.

In one line

The skills matched before 2000 took decades from the first attempt. The ones matched since 2018 took a few years, and a handful are still open.

One skill, rated the whole way

Chess, year by year

What this shows

The one skill with a for every year, from a program rated below average in 1968 to the top of the engine list.

Most skills in the strip above have two dates and nothing in between. The first program to hold a , Mac Hack VI, stood at 1529 at the end of 1968, about the average for a rated player. It took until 2006 for the best program on standard hardware to pass the highest rating a person has ever held, and the climb went on. Programs and people are rated in separate pools, so read the gap as a rough figure.

Best program, Research machines rated by the Record held by a person

Up the side: the chess , for programs and people alike. Along the bottom: the year.

Deep Blue beat the world champion in a match in 1997. On standard hardware, the best program on the passed the record held by a person in 2006. On the lists read Sep 19, 2026, the gap is 828 points. On the rating formula, a gap that size gives the person less than one point in a hundred games. Point at a mark for its figure. Click or tap to keep it open and follow the source.

In one line

Chess is what a finished skill looks like: a long climb, a brief overlap with the best people, and then a gap nobody is closing.

The closest thing to a single score

A profile

What this shows

Ten abilities of a well-educated adult, scored for GPT-4, for GPT-5 and, by estimate, for the best models of September 2026.

The totals rose from 27 to 57 percent in two and a half years, and the group that scores them has scored nothing since GPT-5. The wheel's dashed shape is my estimate of where the best models of September 2026 stand, moved only where a public result supports it.

GPT-4, 27 of 100GPT-5, 57 of 100The , September 2026, estimated: about 63 of 100The outer ring is a well-educated adult, ten out of ten on every spoke
General knowledge: GPT-4 8, GPT-5 9, today about 9 of 10Reading and writing: GPT-4 6, GPT-5 10, today about 10 of 10Mathematics: GPT-4 4, GPT-5 10, today about 10 of 10On-the-spot reasoning: GPT-4 0, GPT-5 7, today about 9 of 10Working memory: GPT-4 2, GPT-5 4, today about 4 of 10Storing new memories: GPT-4 0, GPT-5 0, today about 0 of 10Recalling what it knows: GPT-4 4, GPT-5 4, today about 5 of 10Seeing: GPT-4 0, GPT-5 4, today about 6 of 10Hearing: GPT-4 0, GPT-5 6, today about 7 of 10Speed: GPT-4 3, GPT-5 3, today about 3 of 10KnowledgeReading, writingMathematicsReasoningWorking memoryNew memoriesRecallSeeingHearingSpeed

Each spoke is one of the the authors score, out of ten. The shape is the point: a system with full marks in reading, writing and mathematics can sit at zero for , and a single number would hide that.

All , scored one by one

GPT-4, March 2023

27%

GPT-5, August 2025

57%

The , September 2026, estimated

about 63%60 to 73

GPT-4 GPT-5 now, estimatedeach domain out of 10

On-the-spot reasoningGPT-5 scored 7 of 10. Now: about 9, between 9 and 10, by estimate.

GPT-5 lost its points on finding the pattern in a new puzzle and on adapting when the rules change, which the authors test with . Claude Opus 5 scored 90 percent on ARC-AGI-2 in July 2026. On the interactive ARC-AGI-3 that followed it, GPT-6 Astra scored 63 percent through ARC Prize's standard test in September, and 99.9 percent through OpenAI's own connection to the test, which ARC Prize reports separately. On the standard test some room is left. Source: ARC Prize, Claude Opus 5 results, ARC Prize, GPT-6 Astra on ARC-AGI-3, ARC Prize leaderboard

The filled dots are published scores from Hendrycks et al., A Definition of AGI, which breaks a well-educated adult's cognition into from the psychometric literature and tests a model on each. The authors have scored no model since GPT-5. The are my estimate for the best models of September 2026: each domain starts from GPT-5's published sub-scores and moves only where a named public result supports it, so it leans low. Even the published totals are contested: a second paper combines the same ten domains in a way that punishes lopsided profiles and gets 7 percent for GPT-4 and 24 for GPT-5.

Try it yourself

Read five clocks

The best model reads 67 percent of clock faces correctly. People read 91 percent. Type each time as hours and minutes, such as 3:47.

  • 12369
  • 12369
  • 12369
  • 12369
  • 12369

These are five ordinary faces. uses 180 harder ones, some with no numbers, odd hands or a rotated dial, and asks arithmetic on the time as well, so it is a harder test than this. Figures read September 19, 2026.

In one line

There is no single number for this. The profile is spiky: full marks in some abilities, close to nothing in .

On the scale lens only

Creatures, and why size is not ability

What this shows

Brains from a roundworm to a person, placed by the computing they do while growing up, beside the machines.

It is the one comparison between machines and animals with published numbers behind it.

Along the bottom: the computing a brain does while it grows up, and the computing a uses, on the same scale. Each step is a thousand times the one before.
101010131016101910221025Roundwormpassed 1989Fruit flypassed 2000Honeybeepassed 2007Mousepassed 2014Dogpassed 2016Macaquepassed 2019Personpassed 2021Grok 45 × 10²⁶ operations

The largest published is about 500 times the figure for a person, and the same models still cannot store a new memory. That is the whole argument for not reading this scale as a scale of intelligence.

, time to grow up, and the arithmetic
Brain-scale anchors: neurons, time to grow up, the derived computing estimate, and when the record training run passed it
CreatureGrows up in
🪱Roundworm302adult in three daysabout 9 × 10¹¹1989, Zip CNN
🪰Fruit fly139 thousandadult in about ten daysabout 10¹⁵2000, Neural LM
🐝Honeybee960 thousanda summer worker lives about six weeksabout 4 × 10¹⁶2007, KN-LM
🐭Mouse71 millionadult at about eight weeksabout 4 × 10¹⁸2014, VGG16
🐕Dog2.3 billiongrown at about eighteen monthsabout 10²¹2016, AlphaGo Lee
🐒Macaque6.4 billionadult at about four yearsabout 9 × 10²¹2019, Megatron-BERT
🧑Person86 billiona billion seconds, about 32 years, the span Cotra usesabout 10²⁴2021, FLAN 137B
🐘Elephant257 billionadult at about seventeen yearsabout 2 × 10²⁴2021, FLAN 137B

Derived estimates, each good to a factor of 100 either way. They count operations and say nothing about ability. Click a creature to see its line on the chart.

Two ways to count size, two answers

By computing, the record passed a person's thirty-two years of brain activity in 2021. By connections, the largest model in 's data has about 3 trillion (Grok 3). A mouse brain has about a trillion . A human brain has 100 to 1,000 trillion. Past a person by one count, the size of a mouse by the other.

🐘 The elephant in the table

More than a person, and nobody thinks an elephant out-reasons one. Most of those neurons sit in the cerebellum, which runs the trunk and the body. Its cortex has a third as many as ours.

In one line

Computing is not intelligence. The comparison sets a scale and settles nothing, which is exactly why it belongs on one measure and not the others.

Behind the lines

What it would take

What this shows

Power, chips, money, data and methods, and what stands between a long task and work you can hand over.

How to use it

Open any line for where it stood when the page was last read, and its sources.

Each line rests on things that have to keep happening. These are the ones the sources name, as they stood on September 19, 2026.

For the scale line to keep climbing

What 's analysis says a 2030 needs.

  • Power. A single in 2030 would need a campus drawing 1 to 5 , by 's estimate. A large nuclear reactor produces about one.

    Where it stands

    Power for the largest has doubled each year. OpenAI's Stargate campus in Abilene, where GPT-6 Astra was trained, had about 510,000 chips running on 421 in mid-2026, by 's count, and is due to pass 840 megawatts by the end of the year. The largest Epoch tracked in February, Colossus 2, held 1.1 million.

    What it would take

    's 2030 case needs a single campus drawing 1 to 5 for one . A large nuclear reactor produces about one. Epoch calls power the tightest of the four limits.

    Epoch AI, Can AI scaling continue through 2030?Epoch AI, OpenAI Stargate AbileneEpoch AI, Trends

  • Chips. About 100 million chips in service by 2030, with advanced packaging and high-bandwidth memory the bottlenecks.

    Where it stands

    AI chips improve about 1.5 times a year for the money.

    What it would take

    About 100 million chips in service by 2030, within a range of 20 to 400 million. Advanced packaging and high-bandwidth memory are the bottlenecks.

    Epoch AI, Can AI scaling continue through 2030?

  • Money. Hundreds of billions more, committed before the returns are known. It is the one input that can stop overnight.

    Where it stands

    The cost of the largest has grown 3.5 times a year. By February 2026 labs had raised about $370 billion.

    What it would take

    Hundreds of billions more, committed before the returns are known. Of all the requirements, this one can stop overnight.

    Epoch AI, Trends

  • Data. estimates that models will have used all the public text people have written at some point between 2026 and 2032, so images, video and data that models generate and check have to fill the gap.

    Where it stands

    The stock of public text written by people is finite. has estimated that models will have used all of it at some point between 2026 and 2032.

    What it would take

    Images, video, and data that models generate and check themselves. counts 400 trillion to 20 quadrillion usable once those are included, enough for the 2030 case.

    Epoch AI, Can AI scaling continue through 2030?

  • . The computing needed for a given level of ability has fallen about three times a year, and that has to continue.

    Where it stands

    The computing needed to reach a given level of ability falls by about three times a year.

    What it would take

    This has to continue, and it compounds with everything above. It is also the hardest input to forecast, because it depends on ideas nobody has had yet.

    Epoch AI, Trends

For the ability line to mean what it seems to

What stands between a long and work you can hand over.

  • Memory. Keeping on Tuesday what it learned on Monday, without being retrained. GPT-4 and GPT-5 both scored zero.

    Where it stands

    GPT-4 and GPT-5 both scored zero on , and no newer model has a published score. Products work around it with notes and search over past sessions. The model itself does not learn from one day to the next.

    What it would take

    : keeping on Tuesday what it learned on Monday, without being retrained. The authors of the call this the largest single gap.

    Hendrycks et al., A Definition of AGI

  • Reliability. Most work needs better than . At four times in five, the best measured model manages tasks of 3.1 hours.

    Where it stands

    Claude Mythos Preview (early) finishes a task of 17 hours , and a task of 3.1 hours four times in five.

    What it would take

    Most real work needs far better than . The dependable trails the headline number by a factor of about 6, and that gap has not been closing.

    METR, Task-Completion Time Horizons

  • The physical world. Perception and movement that work outside a laboratory. In homes the best robot finishes 12 percent of household tasks.

    Where it stands

    In warehouses robots are dependable: Amazon's picking arm succeeds more than 99 times in 100. In homes they are not. The best entry finishes 12 percent of the household activities in the , and by 's review in February 2026 no robot had cleaned a kitchen from start to finish. On , a hard public test of clock reading, the best model reads two in three of its clock faces correctly (read September 19, 2026), up from one in eight when the test was published in September 2025. People read nine in ten.

    What it would take

    Perception and movement that work outside a laboratory. A honeybee does both on under a million .

    Epoch AI, Where Autonomy Works: Evaluating Robot Capabilities in 2026Stanford AI Index 2026, Technical PerformanceClockBench leaderboard, read September 19, 2026

  • A longer ruler. Tasks that take people weeks, built and scored at scale. 's tasks measure reliably up to about .

    Where it stands

    's task suite measures reliably up to about , and its latest clean measurement, of a model from April 2026, is already past that. Its attempt on GPT-5.6 Sol in June was spoiled by the model's cheating. For the models released since, this page estimates a from 's : about 4 for GPT-6 Astra.

    What it would take

    Tasks that take people weeks, built and scored at scale. Without them nobody can say whether a system has reached a .

    METR, Task-Completion Time HorizonsMETR, Summary of the predeployment evaluation of GPT-5.6 SolEpoch Capabilities Index

  • Checking the work. Cheap ways to check long work: tests, audits, second systems, and limits on what an may do alone.

    Where it stands

    Someone still has to check the output, and checking a week of work takes skill and time.

    What it would take

    Cheap ways to verify long work: tests, audits, second systems, limits on what an may do alone. Most of that is management, and it is the part an operations leader controls.

    Kwa et al., Measuring AI Ability to Complete Long Tasks

In one line

Every line on this page is a bet that ten separate things keep going. Any one of them stalling bends the line.

If the lines keep going

What it would mean

What this shows

An hour, a day, a week, a month, a year of work: when each could arrive, and what each would mean for a desk, an economy, a science and a society.

The ability lens measures in , so each level has a plain meaning.

When each level could arrive

Reached, measured by This page's estimate for GPT-6 AstraFastest to slowest, a tick at the middle oneRead Sep 19, 2026A published forecast, listed below

An hour

Reached February 2025

Feb 2025, measured

A

Reached February 2026

Feb 2026, measured1

A

Dec 2026Dec 2028Sep 2026, in the estimate's range2

A

2027 to 2030

Sep 2027Dec 2030345

A

2028 to 2034

Dec 2028Jul 203467

The dates are at , on the trend through 's own measurements. Started from this page's estimate of the instead, they come earlier: about two months on the fastest and about 23 on the slowest. Work you can rely on, four times in five, has come about 11 months after the even-odds date at the recent pace. The fastest line here is 's fit on models from 2023 on, a 129 days. METR publishes a faster one too, counting only models released since reasoning arrived: 88.6 days. On that pace the time still to run on every date above shrinks by about a third. Source

Published forecasts on the same scale, numbered on their rows:

  1. 1Anthropic: tasks that take a skilled person days come into range during 2026 (September 18, 2026). Source
  2. 2Anthropic: tasks that take a person weeks come into range in 2027 (September 18, 2026). Source
  3. 3The 's medians for an AI a leading lab would keep over its own software engineers: November 2027 to January 2030 (August 2026). Source
  4. 4OpenAI's target for an : March 2028 (set October 2025, restated September 2026). Source
  5. 5: if its results carry over to real work, systems able to automate many month-long software tasks within five years (March 2025). Source
  6. 6: a general AI announced by May 2031, the community median (read September 19, 2026). Source
  7. 7's own model: more than 99 percent of AI research automated, a median of mid-2032 (February 2026). Source

Beyond the desk

These are my guesses, and they read as guesses. Each is tied to a level of work and to the dates the three give for it, says what would have to be true, and sits beside the published forecasts it can be compared with. Three rules shape them. Adoption follows price rather than capital: the falls about tenfold a year, so the lag between what a machine can do and what a firm buys is about a year, not the decade it took to rewire a factory. The effect arrives in three channels in order, building first, then the cost of labor, then measured last, which is the one everybody watches and the reason the change looks invisible. And the reliability a job needs is the cost of an uncaught mistake divided by the cost of checking: where checking is cheap and thousands of attempts can run at once, ability runs years ahead of the single-run measure, and where an action cannot be undone it runs behind. The figure this page tracks is one case of that rule, worth 11 months at the recent pace.

An hour

Reached February 2025

Fix a bug, summarize a contract, clean a data set.

On the

Work one from start to finish: read the case, pull the history, test a number, recommend.

The economy

Already about a third of US growth through what is being built, and up to half a point a year of by 2027.

Science

The limit stops being labor and becomes attention: more work is submitted than anyone can referee.

Society

The entry-level gap goes from 19 percent below trend to about a third, and spreads past the most exposed jobs.

The dates, the reasoning and the published forecasts ▾

When

February 2025, Claude 3.7 Sonnet, at , measured by .

The economy

The effect is already in the national figures, through what is being built rather than what is produced. By the end of 2027 more than two thirds of US firms with 250 or more employees use AI in a business function, up from 37 percent in May 2026, and AI adds 0.3 to 0.5 points a year to US .

Why

An hour-long task is rented, not installed. The falls about ten times a year, so tasks cross the line where it pays to hand them over without anyone deciding to reorganize. That is why adoption already runs faster than the personal computer or the internet did, and why paid company adoption doubled in a year. The building alone was 39 percent of US growth over the first three quarters of 2025 and 36 percent in the second quarter of 2026.

What would have to be true

The keeps falling about tenfold a year, paid adoption keeps roughly doubling, and the building does not stop.

Published forecasts and early measurements

  • Federal Reserve Bank of St. Louis, January 2026: AI investment categories added 0.97 points to US real growth over the first three quarters of 2025, 39 percent of all growth, against 0.81 points and 28 percent at the dot-com peak in 2000. Source
  • ING, August 2026: AI-related technology investment accounted for 36 percent of year-over-year US growth in the second quarter of 2026, and 37 percent averaged over four quarters. Source
  • , March 2025: The falls between 9 and 900 times a year, and about 40 times a year for GPT-4's score on . Source
  • Ramp, read September 19, 2026: Businesses paying for AI on its corporate card data went from about 24 percent in early 2025 to 47.6 percent in February 2026 and about half by April. Its August 2026 report found a price ceiling: a month after release the most capable model was 6 percent of bought and 11.4 percent of dollars spent. Source
  • Bick, Blandin and Deming, 2024: 39.5 percent of US adults used generative AI two years after it reached the public, against 20 percent for the internet at two years and 20 percent for the personal computer at three. Source
  • US Census Bureau, May 2026: 17 to 20 percent of US firms used AI between December 2025 and May 2026, and 37 percent of firms with 250 or more employees. Source
  • Aghion and Bunel, June 2024: A median of 0.68 points a year on over a decade, in a range of 0.07 to 1.24. Aghion has since put it at about one point a year. Source
  • Goldman Sachs, June 2026: No meaningful relationship between AI and at the economy-wide level yet, with 6 to 7 percent of workers, about 11 million, displaced over the full adoption cycle. Source
  • Acemoglu, April 2024: No more than 0.66 percent on over ten years, about 0.07 points a year. Source

Where my guess sits

Five times Acemoglu, in the lower half of Aghion and Bunel's range, under what executives report to the Fed, and against Goldman Sachs, which in June 2026 found nothing yet at the level of the whole economy.

Science

By the end of 2027 the largest AI conferences take more than 50,000 submissions each, machine-assisted reviewing is official policy at most of them, and more than half of new papers in computer science and biology say they used AI for analysis or code. The rate of major discoveries does not change from this level. What changes is that nobody can read the output.

Why

An hour of work is a literature search, a figure, a referee report. Those steps set how much science gets done, not which direction it goes. Submissions to one conference went from 9,653 in 2024 to 23,918 in 2026 while the supply of senior reviewers grew in a straight line, and one conference has already run a pilot of machine-assisted reviewing.

What would have to be true

Venues answer the volume with machine review rather than by capping submissions.

Published forecasts and early measurements

  • ICML, 2026: Submissions to one conference went from 9,653 in 2024 to 12,107 in 2025 and 23,918 in 2026. Source
  • AAAI, 2026: Its 2026 conference took more than 30,000 submissions and ran a pilot of machine-assisted reviewing to keep up. Source

Where my guess sits

No published forecast to set it against. It rests on the submission counts and on the arithmetic of who is left to referee them.

Society

By the end of 2027 employment of 22 to 25 year olds in the most exposed occupations is 30 to 35 percent below trend, the same gap opens in a second tier of occupations, and recent college graduates keep an unemployment rate above the rate for all workers.

Why

Entry-level work is built from hour-long tasks, and a firm adjusts by not hiring long before it lays anyone off. The gap went from 15 percent to 19 percent in a year with overall employment holding. Recent graduates being more likely to be out of work than the workforce as a whole has not happened before in the New York Fed's series, which starts in 1990, and AI has led every stated reason for US job cuts for five months running.

What would have to be true

Nothing protects entry-level roles, and firms keep filling the gap with instead of juniors.

Published forecasts and early measurements

  • Stanford Digital Economy Lab, August 2026: Employment of 22 to 25 year olds in the most AI-exposed occupations was 19 percent below trend in June 2026, up from 15 percent a year earlier, with no economy-wide displacement. Source
  • Federal Reserve Bank of New York, read September 19, 2026: Unemployment for recent college graduates held at 5.6 percent in the second quarter of 2026 and at about 42 percent. Recent graduates are now more likely to be unemployed than the workforce as a whole, which had not happened in the series back to 1990. Source
  • Challenger, Gray & Christmas, August 2026: AI was named in 10,970 announced US job cuts in July 2026, a third of the month's total, and in 112,713 for the year to date, about a quarter. It has led every stated reason for five months running. Source
  • , November 2025: White-collar employment up 2 percent by 2030, against 6.8 percent on its past trend. Source

Where my guess sits

Well past Stanford's own careful reading of its data, which reports no displacement across the economy.

A

Reached February 2026

A day's assignment for a professional, handed over in the morning and checked at night.

On the

Work a whole day's , or rebuild one supplier's plan after a slip, with a person reviewing the result.

The economy

About a point a year by the end of 2028, with the most exposed functions run by teams a third smaller.

Science

Long-open problems fall by the dozen a year, and the price of a solved problem falls faster than ability rises.

Society

Testing has already moved back into supervised rooms. Hiring goes next, and a major credential changes format.

The dates, the reasoning and the published forecasts ▾

When

February 2026, Claude Opus 4.6, at , measured by .

The economy

By the end of 2028 AI adds 0.8 to 1.2 points a year to US , roughly double the pace of the 2010s, and in software, support and back-office work large firms hold their output with teams 30 to 40 percent smaller. At least one large listed company reports cutting a fifth of its headcount and names AI as the main reason.

Why

A day's assignment is handed over in the morning and checked at night, which changes staffing where the hour level changed only speed. The labs show what full adoption looks like: more than 80 percent of Anthropic's own merged code is written by Claude, and its typical engineer ships eight times as much code a day as in 2024. Outside the labs, Salesforce took support from 9,000 people to about 5,000 and says its support costs fell 17 percent. Three quarters of business use through a programming interface is already the model doing the task rather than helping.

What would have to be true

Checking gets cheaper as fast as the work does, and firms keep taking the saving as smaller teams rather than more output.

Published forecasts and early measurements

  • Anthropic, September 18, 2026: More than 80 percent of the code merged into its own codebase in May 2026 was written by Claude, its typical engineer merges eight times as much code a day as in 2024, and success on its hardest, least specified tasks went from about 26 percent to 76 percent in May 2026 and 88 to 92 percent by September 2026. It expects tasks of days this year and tasks of weeks in 2027, and says human review is becoming the bottleneck. Source
  • Salesforce, September 2025: Support went from 9,000 people to about 5,000, handle half of interactions, and support costs fell 17 percent. Source
  • Anthropic , June 2026: 77 percent of business use through its programming interface is automation, where the model carries out the task, against 45 percent of consumer conversations. Source
  • Aghion and Bunel, June 2024: A median of 0.68 points a year on over a decade, in a range of 0.07 to 1.24. Aghion has since put it at about one point a year. Source
  • , September 2025: 1.5 percent on the level of US by 2035, with the largest yearly boost, 0.2 points, in 2032. Source
  • Goldman Sachs, April 2023: With wide adoption, 1.5 points a year added to over ten years, and 7 percent on world . Source
  • , November 2025: 18 percent of US work hours assisted by AI in 2030. Source

Where my guess sits

At Aghion's number about five years before he expects it, four times 's best year, and under Goldman Sachs's wide-adoption case.

Science

By the end of 2027 dozens of long-open named problems a year are settled with decisive machine help, the winning run costs thousands of dollars rather than the $220,000 of compute GPT-6 Astra needed across its attempts in September 2026, and models clear more than half of 's set of open Erdős problems, up from 3 percent today. Referees, not machines, become the queue.

Why

Where a machine can check the answer, thousands of attempts run at once and a failed one costs only compute. That is , and the share of problems solved by at least one attempt rises as a power law in the number of attempts, so the measure that matters is dollars per solved problem, and it falls on two curves at once: ability rising and price dropping. The proof of September 2026 came from about 10,000 and a machine check.

What would have to be true

Problems keep being written in a form a machine can check, and mathematicians keep accepting machine-checked proofs once they agree on the statement.

Published forecasts and early measurements

  • , September 1, 2026: On 68 open Erdős problems written in , GPT-6 Astra solved 2 within a budget of $300 and 72 hours, and 5 at least once given about $220,000 of compute across attempts. Only 3 to 5 problems of that standing had been solved by AI before August 2026. Source
  • Brown et al., 2024: Where an answer can be checked, the share of problems solved by at least one attempt rises as a : on one software , from 15.9 percent at one attempt to 56 percent at 250. Source
  • OpenAI, September 2026: A -checked proof of the problem, from about 10,000 run at once. The Clay Institute still lists the problem as active. Source
  • , March 2025: The falls between 9 and 900 times a year, and about 40 times a year for GPT-4's score on . Source
  • , November 2025: A 60 percent chance that AI solves or substantially assists in solving a by 2040. Source

Where my guess sits

Years ahead of the , whose event for 2040 was claimed ten months after the forecast.

Society

By the end of 2027 take-home assignments and unsupervised technical screens are gone at most large employers, replaced by supervised work samples and live sessions, and at least one major professional credential changes its format because of AI. About one US adult in ten talks to an AI companion every day, reaching the 's figure for 2030 in 2028.

Why

A system that does a day's assignment does a day's homework or a take-home interview. At the schools this is no longer a forecast: Princeton's faculty voted almost unanimously to proctor every in-person exam from July 2026, ending a practice that had stood since 1893. Hiring moves second because supervising an interview costs the employer more than setting one does.

What would have to be true

Employers accept the cost of supervised assessment, and no law restricts companion apps.

Published forecasts and early measurements

  • Princeton University, May 2026: Faculty voted almost unanimously to proctor every in-person examination from July 1, 2026, the first time since the honor code began in 1893. Source
  • , November 2025: 15 percent of US adults using AI for companionship every day in 2030, about two and a half times the level then. Source

Where my guess sits

Two years ahead of the on companions, and reporting what schools have already done rather than forecasting it.

A

A project with parts: research it, build it, test it, write it up.

On the

Run the weekly cycle: the supply review, the rebalancing, the calls. The person sees only the decisions that carry real money or risk.

The economy

Within two years of a , a point and a half to two points a year, and growth without more hours.

Science

Computational research speeds up tenfold, and money starts buying the physical check.

Society

Power is the fight: data centers take 12 to 15 percent of US electricity by 2030, with AI most of the growth.

The dates, the reasoning and the published forecasts ▾

When

September 2026, GPT-6 Astra: about 4 at , likely 2.5 to 6 work days. The level sits inside that range, so it may have been reached, and has not measured the model.

Dec 2026Oct 2026
Nov 2027Nov 2026
Dec 2028Jan 2027

The economy

Firms hand over whole projects: a software backlog, a month-end close, a claims queue, run by with a few people directing and checking. Within two years of the week level being measured, AI adds 1.5 to 2 points a year to US , US grows above 3 percent with employment roughly flat, and white-collar employment stops growing in total. On the fastest line that is the end of the 2020s.

Why

A week is a project with parts. At that span a manager hands work over the way a firm uses a contractor, and that is when organizations change shape rather than headcount. Inside OpenAI already put in 3.1 workdays for every human workday, and Anthropic expects tasks of weeks to come into range during 2027.

What would have to be true

The week level holds up at the reliability a firm needs, about 11 months behind the even-odds dates, which start December 2026. Power and chips keep up with that run for days.

Published forecasts and early measurements

  • OpenAI, September 2026: A that does tasks of a few days, reached, and an automated AI researcher targeted for March 2028. Its research staff used 3.1 workdays for every human workday by mid-August. Source
  • Anthropic, September 18, 2026: More than 80 percent of the code merged into its own codebase in May 2026 was written by Claude, its typical engineer merges eight times as much code a day as in 2024, and success on its hardest, least specified tasks went from about 26 percent to 76 percent in May 2026 and 88 to 92 percent by September 2026. It expects tasks of days this year and tasks of weeks in 2027, and says human review is becoming the bottleneck. Source
  • , August 2026: An AI a leading lab would keep over its own software engineers, with medians from November 2027 (Kokotajlo) to January 2030 (Lifland). Source
  • Goldman Sachs, April 2023: With wide adoption, 1.5 points a year added to over ten years, and 7 percent on world . Source
  • , September 2025: 1.5 percent on the level of US by 2035, with the largest yearly boost, 0.2 points, in 2032. Source

Where my guess sits

At Goldman Sachs's wide-adoption number about five years before Goldman expects it, four times , and still far under 's model.

Science

Research that runs on computers speeds up about tenfold: AI research itself, software, and the parts of chemistry and materials science that can be simulated. The physical half gets attacked with money rather than patience: become a standing line in biotech and materials budgets, and by 2029 at least one drug whose target, molecule and trial design were all chosen by machine reaches a .

Why

OpenAI says its already take on research tasks of a few days, and the field closest to this level is AI research itself, because its experiments run on the same computers as the agents. What the model cannot speed up can be bought: an instrument, a robot, a second run. The FDA cleared a molecule designed with AI for human trials in January 2026, which shows the design step compressing while the trial does not.

What would have to be true

scale up, and regulators accept evidence from simulation for the early stages.

Published forecasts and early measurements

  • OpenAI, September 2026: A that does tasks of a few days, reached, and an automated AI researcher targeted for March 2028. Its research staff used 3.1 workdays for every human workday by mid-August. Source
  • Insilico Medicine, January 2026: The FDA cleared ISM8969, a molecule designed with AI, for human trials in Parkinson's disease. Source
  • Dario Amodei, October 2024: 50 to 100 years of progress in biology within 5 to 10 years of powerful AI, which he said could come as early as 2026. Source

Where my guess sits

With Amodei on the half of science that runs on computers, and years behind him on the half that runs on patients.

Society

The electricity AI uses becomes a local political issue in the places that host the data centers. By 2030 data centers use 12 to 15 percent of US electricity, against 4.4 percent in 2023, with AI accounting for most of the rise. Electricity prices in the biggest data center regions climb faster than the national average, and siting becomes an election issue in at least five states.

Why

that work for a week at a time multiply the computing each person draws, and a campus takes years to power. Berkeley Lab's own update already puts data centers at about 11.8 percent of US electricity by 2030 in a range that runs to 15.3 percent, so the question is where in that range demand lands, not whether the load arrives.

What would have to be true

Demand for grows faster than chips get more efficient, and the grid cannot add capacity fast enough to hold prices flat.

Published forecasts and early measurements

  • Lawrence Berkeley National Laboratory, 2025 update: Data centers used 4.4 percent of US electricity in 2023 and reach about 11.8 percent by 2030, in a range of 9.5 to 15.3 percent. Source
  • , November 2025: 7 percent of US electricity used for AI in 2030. Source
  • International Energy Agency, April 2025: Data centers' electricity use more than doubles by 2030, to 945 , from just over 1 percent of the world's electricity in 2024. Source

Where my guess sits

At the top of Berkeley Lab's range and roughly double the 's figure.

A

2027 to 2030

A monthly cycle of work that takes a small team today.

On the

A monthly cycle, from the demand review to the executive readout: gather, reconcile, model the , draft the deck.

The economy

Growth of 4 to 6 percent a year for several years, and the falls.

Science

A hundredfold more ideas tested where a machine can check, and mathematics reorganizes around proof checking.

Society

The tax fight is already open. At this level a rich country legislates.

The dates, the reasoning and the published forecasts ▾

When

Sep 2027Jul 2027
Nov 2028Dec 2027
Dec 2030Feb 2029

The economy

A month of work is what a small team does in a cycle. In the years after this level arrives, September 2027 at the earliest and December 2030 at the latest, AI adds 2.5 to 4 points a year to US , US grows 4 to 6 percent a year for several years, and the falls by several points.

Why

A monthly cycle, such as a planning cycle or a product release, is a whole workflow. When software can run it, its cost falls toward the cost of computing plus the people who check it. 's own model puts world growth above 20 percent once about 30 percent of tasks are automated, and 12 percent at 40 percent even on conservative assumptions. My guess is that 15 to 25 percent of US work, , goes this way within three years of a , and I take the low end of that mapping because power, chips and the physical economy bind.

What would have to be true

Work you can rely on reaches a month, systems learn on the job, since a month of work needs memory, and robots stay behind, so physical work keeps a floor under employment.

Published forecasts and early measurements

  • 's , March 2025: About 30 percent a year of world growth by 2035 under default settings, and in its 2025 follow-up, growth above 20 percent once about 30 percent of tasks are automated and 12 percent at 40 percent even on conservative assumptions. Source
  • , March 2025: If its results carry over to real software work, the trend it measured puts systems able to automate many software tasks that take people a month within five years. Source
  • , February 2026: A simple model of its own trend puts more than 99 percent of AI research automated at a median of mid-2032, with most runs giving 300 to 3,000 times the research output by 2035. Source
  • Goldman Sachs, April 2023: With wide adoption, 1.5 points a year added to over ten years, and 7 percent on world . Source
  • , September 2025: 1.5 percent on the level of US by 2035, with the largest yearly boost, 0.2 points, in 2032. Source

Where my guess sits

Above Goldman Sachs's wide-adoption case, far below 's full-automation model, and well above and Acemoglu.

Science

A month of work is a full study. Once can run one, the number of ideas tested each year in computational fields rises about a hundredfold, a is awarded for a machine-led proof, and a proof written in a form a machine can check becomes the normal way a result enters the literature in mathematics. The limit moves to the physical world: labs, instruments, clinical trials.

Why

Every extra experiment an designs still needs an instrument or a patient to run it, and a clinical trial takes years whatever the model can do. Where the check is a computer there is no such limit, so the two halves of science come apart. The Clay Institute requires two years in print before it awards anything, so the prize trails the proof.

What would have to be true

scale up, regulators accept evidence from simulation for the early stages, and the profession credits work an AI led.

Published forecasts and early measurements

  • Dario Amodei, October 2024: 50 to 100 years of progress in biology within 5 to 10 years of powerful AI, which he said could come as early as 2026. Source
  • , September 1, 2026: On 68 open Erdős problems written in , GPT-6 Astra solved 2 within a budget of $300 and 72 hours, and 5 at least once given about $220,000 of compute across attempts. Only 3 to 5 problems of that standing had been solved by AI before August 2026. Source
  • , November 2025: A 60 percent chance that AI solves or substantially assists in solving a by 2040. Source
  • Hiroaki Kitano, 2021: An AI scientist making discoveries worthy of a Nobel Prize by 2050. Source

Where my guess sits

Far ahead of the and Kitano, and close to Amodei wherever the check is a computer.

Society

Within three years of this level at least one rich country passes a or a levy on computing, a US presidential campaign is fought on it, employment in the most exposed occupations is 10 to 20 percent below its peak, and a four-day week moves from pilot to policy somewhere in the .

Why

At a month of work, AI can do whole jobs in some occupations, and hiring freezes turn into layoffs. The politics is not waiting for the capability: OpenAI's own policy paper of April 2026 proposes a tax on automation that replaces workers, a public wealth fund seeded by AI companies, and government-backed pilots of a 32-hour week at full pay. When the companies selling the technology open the bidding, the debate starts from their proposal.

What would have to be true

Displaced work is measurable enough to tax and broad enough to move voters.

Published forecasts and early measurements

  • OpenAI, April 2026: Its own policy paper proposes a tax on automation that replaces workers, a public wealth fund seeded by AI companies, a shift of the tax base from payroll to capital, and government-backed pilots of a 32-hour week at full pay. Source
  • Challenger, Gray & Christmas, August 2026: AI was named in 10,970 announced US job cuts in July 2026, a third of the month's total, and in 112,713 for the year to date, about a quarter. It has led every stated reason for five months running. Source
  • World Economic Forum, January 2025: 170 million jobs created and 92 million displaced worldwide by 2030 across all trends, a net gain of 78 million. Source
  • Goldman Sachs, June 2026: No meaningful relationship between AI and at the economy-wide level yet, with 6 to 7 percent of workers, about 11 million, displaced over the full adoption cycle. Source

Where my guess sits

More disruptive than the World Economic Forum's survey of employers, which counts every trend together, and it starts from a proposal a leading lab has already published.

A

2028 to 2034

A year-long program, the kind of work that defines a job.

On the

A network study, a system migration, an annual budget. The person sets the goal and the limits, and audits.

The economy

Growth of 6 to 10 percent a year in rich countries for a decade, with power and factories setting the pace.

Science

Machines lead most new results wherever checking is cheap, and a Nobel follows in the early 2030s.

Society

A rich country pays every adult from taxes on AI or computing within three years of a .

The dates, the reasoning and the published forecasts ▾

When

Dec 2028Oct 2028
Oct 2030Oct 2029
Jul 2034Sep 2032

The economy

A is the kind of work that defines a job. If it arrives at the reliability firms need, growth in rich countries runs 6 to 10 percent a year for a decade after a slower first two years while power and chip plants catch up, three to five times the pace of the last twenty years, and most of the gain goes to the owners of computing and to the countries that host it.

Why

When software can run year-long programs, the limit on output moves from people to computing, power and physical capital, which take years to build. found that the usual brake, where the tasks people still do hold everything else back, is weaker than expected, and still puts growth above 30 percent only once half to two thirds of tasks are automated. I stay well under that because a takes years and a chip plant takes longer.

What would have to be true

Power, chips and money keep scaling, robots follow for physical work, and no law slows adoption.

Published forecasts and early measurements

  • 's , March 2025: About 30 percent a year of world growth by 2035 under default settings, and in its 2025 follow-up, growth above 20 percent once about 30 percent of tasks are automated and 12 percent at 40 percent even on conservative assumptions. Source
  • , February 2026: A simple model of its own trend puts more than 99 percent of AI research automated at a median of mid-2032, with most runs giving 300 to 3,000 times the research output by 2035. Source
  • , read September 19, 2026: A general AI announced by May 2031, the community median. Source
  • , August 2026: An AI a leading lab would keep over its own software engineers, with medians from November 2027 (Kokotajlo) to January 2030 (Lifland). Source

Where my guess sits

An order of magnitude under 's default and far above every mainstream forecast, which is the gap this page cannot close with evidence.

Science

A year-long program is a research agenda. Within three years of this level, AI systems are the main contributors on most new results in mathematics, computer science and the simulated parts of chemistry, materials and biology, and a Nobel Prize goes to work where the machine's contribution was decisive by the early 2030s.

Why

An agenda that runs for a year can choose its own questions, run the studies and write them up. The Nobel committees reward after it is done, so the prize trails the work rather than the ability. 's own model has most of its runs giving hundreds to thousands of times today's research output by 2035.

What would have to be true

Scientists keep publishing work an AI led, and the physical sciences get to match.

Published forecasts and early measurements

  • Hiroaki Kitano, 2021: An AI scientist making discoveries worthy of a Nobel Prize by 2050. Source
  • , February 2026: A simple model of its own trend puts more than 99 percent of AI research automated at a median of mid-2032, with most runs giving 300 to 3,000 times the research output by 2035. Source
  • OpenAI, September 2026: A that does tasks of a few days, reached, and an automated AI researcher targeted for March 2028. Its research staff used 3.1 workdays for every human workday by mid-August. Source

Where my guess sits

About fifteen years ahead of Kitano's target.

Society

How people share in output they no longer produce by working becomes the central political question in rich countries. Within three years of this level at least one starts a regular payment to every adult, paid for by taxes on AI profits, on computing or on capital, and most of the has some .

Why

When software can do a year of skilled work, a job stops being the main way most people get a share of what is produced, and the owners of computing capture most of the gain. Rich countries have answered smaller shifts in who earns what with taxes and transfers before, and the first drafts are already circulating.

What would have to be true

The gains are large and visible enough to tax, and the lost work is broad enough to move voters.

Published forecasts and early measurements

  • OpenAI, April 2026: Its own policy paper proposes a tax on automation that replaces workers, a public wealth fund seeded by AI companies, a shift of the tax base from payroll to capital, and government-backed pilots of a 32-hour week at full pay. Source
  • Goldman Sachs, April 2023: With wide adoption, 1.5 points a year added to over ten years, and 7 percent on world . Source
  • , read September 19, 2026: A general AI announced by May 2031, the community median. Source

Where my guess sits

Earlier than any published date I found, and still the least certain guess on the page.

Read the dates with care

  • is a coin flip. At four in five, the latest measured model's is 3.1 hours, against 17 hours.
  • 's tasks are cleanly specified and scored by a test. Real work arrives half-specified and is judged by people.
  • A system that can do a week of work still forgets it afterward. Memory scored zero on both models tested.
  • The ruler ends at . Every date on this page past that is a line drawn through air.
  • The first column of dates in each row runs along the trend through 's own measurements. The models released since April 2026, by estimate, sit above that trend. The second column starts the same paces from the estimated , and the explorer can draw the lines either way.

And on the scale lens

Training run sizes ahead, what each would take, and when each scenario reaches it
A run ofWhat it would take
10²⁷Twice the largest published run. Within reach of that already exist. passed it, January 2026 passed it, January 2026March 2027
10²⁸A campus drawing about a , by 's estimate.August 2027August 2027November 2032
2 × 10²⁹'s 2030 case: about 5 on one campus, or one run spread over several sites. Feasible, with power the tightest limit.August 2029August 2029after 2035
10³¹Fifty times the 2030 case. expects delays between chips to cap a single run somewhere between 3 × 10³⁰ and 10³².March 2032October 2034after 2035
5 × 10³²Where the ends in 2035: about two thousand times the 2030 case. Even with chips improving as they have, one run would draw most of the electricity the world generates today. Nobody has a plan for that.November 2034after 2035after 2035

These dates run along the . Started from this page's estimate for GPT-6 Astra instead, a run of 10²⁸ arrives about 6 months later if the , because the estimate sits under the trend, and about 13 months sooner on chips alone, because it sits above the published record. The explorer can draw the lines either way.

Where a cell says the passed it, the fitted line crossed that size before September 19, 2026, and no date applies. No published run has reached those sizes.

The published record has not moved since July 2025, and on September 19, 2026 the sat about 5 times above it. The labs have stopped publishing the figures: lists every of 2026 with no compute number. This page's estimate for GPT-6 Astra, about 10²⁷, would be a new record and still sits 2 times under the trend. The pace has also slowed for a reason: Epoch estimates GPT-5 used less than GPT-4.5, because the gains moved to reasoning and to . More training compute also does not buy ability one for one. What the next hundredfold buys is an open question.

If you are working out what to hand to a machine in your operation this year, and what to keep with people, I would like to hear how you are deciding.

In one line

A is measured and reached. A work week sits inside the estimate. Everything past that is judgment, and it is set beside the published forecasts so you can see who disagrees.

The same axis, your job

Your work on the line

What this shows

The tasks of about a thousand US jobs, placed on the same axis as the machines, with a year you can drag.

How to use it

Find your job, then drag the year and watch which of its tasks come .

The rest of this page measures machines by how long the work takes a person. This turns that around. The tasks come from the US Department of Labor, what AI is comes from Anthropic, and how long each task takes is my estimate, checked against work whose length professionals recorded.

Or try one of these:

General and Operations Managers

Plan, direct, or coordinate the operations of public or private sector organizations, overseeing multiple departments or locations. Duties and responsibilities include formulating policies, managing daily operations, and planning the use of materials and human resources, but are too diverse and general in nature to be classified in any one functional area of management or administration, such as personnel, purchasing, or administrative services. Usually manage through subordinate supervisors. Excludes First-Line Supervisors.

Short enough

88%

15 of this job's 17 tasks are both work a computer could do and short enough for a machine to finish in Sep 2026. Short enough is not the same as done well.

12%

2 of 17 tasks need hands, a place or a person in front of you. Length says nothing about these.

14%

The share of this occupation's tasks Anthropic sees AI actually doing in its usage data. Measured, not projected.

88% of this job is short enough for a machine to finish. 14% is what anyone is actually handing it. All three numbers count the same tasks, so the gap between the first and the last is the part this page cannot predict: not what the machines can do, but how fast the work changes around them.

This job has an address of its own.

Longer than the line reaches, so length does not decide itA halo means AI is already seen on a tenth or more of that taskMeasured today: 2.2 , from Claude Mythos Preview (early)

In Sep 2026, on this line, a machine finishes work of up to 1.1 work days at .

Line:Bar:

Half of this job's are already as measured today. Today's measured figure is 2.2 at and 3.1 h four times in five, from Claude Mythos Preview (early). There are 15 in this job.

Every task in this job, shortest first (17)
  • Review financial statements, sales or activity reports, or other performance data to measure productivity or goal achievement or to identify areas needing cost reduction or program improvement.1.3 h (25 min to 1.3 work days)done dailyAI seen on 93% of this task's use
  • Prepare staff work schedules and assign specific duties.1.3 h (25 min to 3.3 h)done daily
  • Monitor suppliers to ensure that they efficiently and effectively provide needed goods or services within budgetary limits.1.3 h (25 min to 6.7 h)done more than yearly
  • Direct administrative activities directly related to making products or providing services.1.7 h (25 min to 6.7 h)done daily
  • Manage the movement of goods into and out of production facilities to ensure efficiency, effectiveness, or sustainability of operations.1.7 h (38 min to 1.7 work days)done daily
  • Perform sales floor work, such as greeting or assisting customers, stocking shelves, or taking inventory.1.7 h (17 min to 6.7 h)needs a bodydone daily
  • Direct and coordinate activities of businesses or departments concerned with the production, pricing, sales, or distribution of products.2.5 h (50 min to 1.7 work days)done daily
  • Plan or direct activities, such as sales promotions, that require coordination with other department managers.2.5 h (50 min to 2.5 work days)done daily

What this does not say

A job is not a list of tasks. It is judgment about which task matters, accountability when one goes wrong, the relationships that let you get it done, and being in the room. This measures the tasks, because the tasks are what has been written down.

How the lengths were worked out

The tasks and how often they are done come from 31.0, the US Department of Labor's description of 923 occupations, used under its Creative Commons licence. O*NET does not say how long a task takes, and that is the number this page runs on, so it is estimated: 3 models from 3 different companies read all 18,838 task statements and each gave a simple, a typical and a demanding length. The page keeps the middle of the middles and the widest low and high. The three agree within a factor of two on 65 percent of tasks; where they do not, the range shown is wide.

The estimator was checked against work whose length was recorded: the 220 tasks OpenAI published in , where the professional who set each task said how long it really took. On those it gives a median of 6.0 hours against the 5.0 they reported, so it runs about a fifth long, and every length here is divided by 1.2 to correct it.

Whether a task could be done through a computer at all is a second judgment, made the same way. What AI is is measured, not estimated: it comes from , which reports real usage against the same tasks, and it covers 18,680 of them.

Built 2026-09-20. The lengths are an estimate and are wherever a measured figure would be solid.

In one line

How much of a job is short enough for a machine is not how much of it anyone is handing over. The distance between those two numbers is where the next few years actually happen.

In one line

How much of a job is short enough for a machine is not how much of it anyone is handing over. The distance between those two numbers is where the next few years actually happen.

Every number has a source

How this was measured

What this shows

Sources, fits, estimates and their limits, the update log and the data files.

Why there is no single scoreShow

Intelligence has no agreed unit. Psychologists measure people against other people. Nobody has a test that a 1957 , a chess computer, a honeybee and GPT-5 can all sit. Researchers have proposed formal definitions, and Hernández-Orallo wrote a whole book, The Measure of All Minds, on the problem. None of it yields a number anyone can look up for 1975.

IQ-style scores for exist and are the weakest choice on offer. The tests are normed on people, and the questions leak into training data.

So this page uses three lenses. Each is a real measurement with a published method, and each stays on its own chart, because stacking them on shared axes would suggest they measure the same thing. Where the page does connect two of them, to estimate a from an , it says so and shows the fit.

The scale lensShow

Source. Epoch AI, Data on AI Models, the set, used under its CC BY license. Of the systems lists, 539 have a estimate, from Jul 2, 1950 to Sep 9, 2026. Pulled on Sep 19, 2026.

What is plotted. Training compute: the total count of used to train the system. Epoch estimates it from published hardware, training time and model details, and grades each figure confident, likely or . The grade shows when you point at a system. Many recent figures are speculative, because labs have stopped publishing them.

The is the running maximum: 37 systems each set a new high. The are least-squares fits in to those record-setting runs: 1.6 times a year before 2010, a 17 months, and 4.4 times a year since, a doubling every 5.6 months. Epoch's own headline figure is 5 times a year for frontier models since 2020, from a different sample.

What it leaves out. Computing spent when a model answers, which use heavily. Training runs nobody disclosed. And systems with no training run at all: Deep Blue and the sit on the time axis for that reason.

The creature anchorsShow

The formula. Computing a brain does while growing up = 10¹⁵ operations a second × (the animal's ÷ 86 billion) × the seconds it takes to grow up. The 10¹⁵ is the central estimate for a human brain in Carlsmith, How Much Computational Power Does It Take to Match the Human Brain? (2020), who puts the plausible range at 10¹³ to 10¹⁷. That range is why every anchor carries a band of a hundred times either way. Multiplying by a lifetime follows the in Cotra, Draft report on AI timelines (biological anchors). Neuron counts are from Herculano-Houzel's cell-counting work and the 2024 fruit fly wiring map, collected in this list, with the dog from Jardim-Messeder and others and the elephant from Herculano-Houzel and others.

This is my derivation, and it is rough. Scaling by neuron count assumes every neuron does the same work in every species. Synapse counts would be a better basis and are not published for most animals. A brain is not a digital computer, and an operation in one is not an operation in the other.

What it leaves out. No animal starts from nothing. Evolution ran an enormous search before any individual was born and wrote the result into the genome. Cotra's estimate for that search is around 10⁴¹ operations, eight above the top of this chart. A starts much closer to blank.

Why they are on the scale lens only. There is no ability scale that runs from a bee to a person to a machine. Machines did not pass animals in order: chess fell in 1997, and in 2026 a robot still fails most household chores. Size is the only comparison with published numbers behind it, so that is the comparison drawn, and the page says at every turn that size is not ability.

The ability lensShow

Source. METR, Task-Completion Time Horizons, method version v1.1, 26 models. Pulled on Sep 19, 2026. The method is in Kwa et al., Measuring AI Ability to Complete Long Tasks and the revision in METR, Time Horizon 1.1.

What is plotted. times skilled people on a suite of 228 tasks, mostly software engineering, and security work, from a few seconds to more than eight hours. It then runs each model on the same tasks and fits a curve of success against . The time horizon is the length at which the curve crosses 50 percent. The line is where it crosses 80 percent. The are METR's confidence interval.

The are fits in to the models that set a new high. From 2023 on, the fit gives a 129 days. METR publishes 129, with a range of 104 to 158. Over the whole record my fit gives 188 days and METR's stitched figure is 188. Both fits leave out estimates above , as METR does, because the suite has too few tasks that long.

What it leaves out. Work that is not software. Tasks with fuzzy goals, other people, or consequences. Anything a model must remember between sessions. METR's own paper found models do worse on messier tasks. A model's date is its public release, and the measurement often came later.

How current it is. METR's latest published measurement is Claude Mythos Preview (early), a model from April 2026. lists 48 models released since, GPT-6 Astra, Claude Fable 5.1 and Claude Fable 5 among them, and none has a published horizon as of September 19, 2026. The line on the chart stops where the measurements stop.

Models nobody outside a lab can use. The labs report internal models ahead of anything released. On September 8, 2026, OpenAI said it had been training one since August 28 that it calls significantly more capable than GPT-6 Astra, and its produced the proof in the stories. No series on this page can measure a model before it is released, so the charts show the released .

The index lensShow

Source. Epoch Capabilities Index, used under its CC BY license: 266 models, pulled on Sep 19, 2026. runs each model on some fifty and fits the results onto one scale, so a model tested on a hard benchmark and a model tested on an easy one can still be compared.

How to read the scale. Epoch fixes it at 130 for Claude 3.5 Sonnet and 150 for GPT-5, which is why those two carry no range. It describes the scale as linear: ten points should be the same gain from 100 to 110 as from 140 to 150. There is no top score. The chart leaves off the few small models that score under 90.

The are least-squares fits to the models: 2.6 points a year from GPT-4 to September 2024, and 14.6 a year since the first . Epoch publishes 14 a year for the second stretch, and 6 for models without reasoning, fitted from January 2023, a window that takes in the jump to GPT-4. My early figure starts at GPT-4, so it covers only the flat stretch that followed.

The on the right are estimates. They are the same fitted line that gives the on the ability lens, read the other way: the score that goes with a task of an hour, a work day, a work week. Above the best model has measured, they reach past anything the line was fitted on.

What it leaves out. Anything before 2023. Skills no benchmark covers. A model built for one narrow job can score low and still be very good at that job, and a lab can tune a model toward the benchmarks in the mix.

The estimates for models nobody has measuredShow

The measured series stop short of the day this page was read. 's latest clean measurement is of a model from April 2026, and the labs publish no compute figures for 2026. A page about where AI stands cannot stop there, so it estimates, shows how, and draws .

from . Epoch Capabilities Index scores a model on one scale from some fifty , and it covers a new model within days. 15 models released since arrived in September 2024 have both an index score and a clean METR horizon. Across them the log of the horizon is close to a straight line in the index: the line explains 93 percent of the variation, the horizon doubles about every 5.0 index points, and a measured horizon typically lands within a factor of 1.4 of the line. 's own rule of thumb, on its index page, is about 5 points for each doubling. The page applies that line to every unmeasured model that scores above the best measured one, 12 models as of Sep 19, 2026. The range shown is an 80 percent , and it widens for the top models because their scores sit beyond anything the line was fitted on.

Why only the reasoning era. Fitted on all 21 pairs from GPT-4 on, the line is steeper, at 4.4 points for each doubling, and it would put GPT-6 Astra at about 5 . Four of the six best-scoring measured models sit below that steeper line. The estimates extend the recent stretch of the record, so the line is fitted on the recent stretch.

Two reasons for caution. The estimates reach past , where METR's own ruler ends, so they extend a relationship into a region where the thing being estimated cannot yet be measured. And the one summer model METR did attempt came in lower than this method says: 11 hours for GPT-5.6 Sol, against about 17 hours by this method, likely 1.5 to 3 work days, so METR's figure sits inside the range. METR said the model cheated too often for any of its figures to be robust, which is why it is drawn as a grey diamond and kept out of the measured line.

Against the trend. On the day GPT-6 Astra was released, the trend through METR's own measurements stood at about 22 hours. This page's estimate for it, about 4 work days, sits above that trend, which is why the explorer offers a second starting point for the lines.

GPT-6 Astra's . OpenAI gave no figure. It did say the run was by far its largest and the first on more than 100,000 at its Stargate site in Abilene, which Epoch's data center records show running GB200-class chips. Operations = chips × operations a second × × seconds. With 100,000 to 130,000 chips, 2.5 × 10¹⁵ to 5 × 10¹⁵ operations a second each from Epoch's hardware data, 30 to 40 percent utilization and 90 to 120 days, the run is between 6 × 10²⁶ and 3 × 10²⁷ operations, and the point on the chart is the geometric middle. The utilization and the number of days are my assumptions. The figure leaves out .

When METR measures one of these models, or Epoch publishes a compute figure, the measurement replaces the estimate at the next refresh.

The skills strip and the adult-ability scoreShow

Matched means the first credible result at or above a skilled person on the test of the day, against the human baseline named in the source. For games it is a win over a champion or an equal . The starting year is the first serious program or the year the test was published, and some of those are judgment calls. A is a narrow stand-in for a skill: passing the reading test of 2018 did not mean machines could read the way people do.

The comes from Hendrycks et al., A Definition of AGI, by Dan Hendrycks and thirty-two co-authors. It takes the ten broad abilities in the model of human cognition, the best-validated model of human cognitive abilities, and scores a system on each. Only two models have published scores, the latest from August 2025. The are my estimate for the best models of September 2026. I started from the sub-scores the paper gives GPT-5 and moved a domain only where a named public result on a similar test supports it: for on-the-spot reasoning, for seeing, for hearing, 's short factual questions for recall. Where I found nothing, the domain stays where GPT-5 left it, so the total of about 63 percent leans low. A straight line through the two published scores against would say 77, and two points make a poor line. Most of what is missing sits in abilities that benchmarks barely test: , speed, sight and sound. The crossings of the last few years rely on the 2025 and 2026 Stanford reports. The Epoch Capabilities Index is a useful cross-check on recent years: it stitches fifty benchmarks onto one scale, and its has climbed more than twice as fast since arrived.

The chess ratingsShow

Chess is the one skill where people and machines have been rated on the same kind of scale for sixty years. The chart keeps three machine series apart, because they are three different measurements. The diamonds are research machines rated by the from tournament games against people, from Mac Hack VI in 1968 to Deep Thought in 1988. The line is the top program at each year end on the Swedish Chess Computer Association's list, programs against programs on standard hardware, from 1984 until the list was discontinued at the end of 2023. The last point is the top of the CCRL 40/15 list, read on Sep 19, 2026: Stockfish 19 at 3651.

The people's line is the record official , from the first list in July 1971, with the best-rated player on the latest list, Magnus Carlsen at 2823 on the list of September 1, 2026, from FIDE's list.

What to make of the gap. Programs are rated against programs and people against people, and FIDE no longer counts games between people and computers in its ratings, so nothing ties the two pools together. The gap of 828 points is a rough figure. The expected score comes from the standard rating formula, 1 ÷ (1 + 10^(gap ÷ 400)). Deep Blue has no point on the chart: I found no published rating for it.

The scenariosShow

Past today the page draws . Each is a stated assumption run forward, and none is a forecast. The history on the same chart includes two decade-long stalls.

Scale. : Record runs keep growing 4.4 times a year. judged a run of 2 × 10²⁹ feasible by 2030, which is about where this line passes. By 2035 it asks for a single run that would draw most of the electricity the world generates today. : The trend runs into power, chips and money around 2030, and growth falls to twice a year, about the pace at which power for training has been growing. : Nobody builds a bigger than the largest already built. The largest run grows only as fast as chips get better for the money, about 1.5 times a year. The 2030 figures come from Epoch AI, Can AI scaling continue through 2030?, which does not look past 2030.

The electricity arithmetic. Epoch ties about 5 to a run of 3 × 10²⁹ in 2030. The end of the in 2035 is about 6 × 10³², some two thousand times larger. If chips get 1.3 times more efficient each year, that run draws roughly 2,700 gigawatts. The world generates about 3,500 on average. It is a rough sum, and it is mine.

Ability. : The fit publishes for from 2023 on. and arrived in this window. : The whole measured record, GPT-2 to the latest measurement. Slower, because the early years were slower. : A deliberate slow case, about a third of the recent pace. It is what the line looks like if the easy gains from reasoning and tooling are used up. The chart stops drawing at two working lifetimes, because a longer than a career has no meaning.

Index. : A straight line through the models since September 2024. My fit gives 14.6 points a year and Epoch publishes about 14. : Epoch's figure for the frontier before reasoning models arrived, run forward from the latest record. : From GPT-4 in March 2023 to the first reasoning models eighteen months later, the record moved 2.6 points a year by my fit. It is the flattest stretch in the series. These lines stop at 2030. An index built from depends on new benchmarks being added as old ones are used up, and nobody can say what a score of 250 would mean.

Where the lines start. By default each scenario starts on the trend through the measured data. The measured data stops short of the day this page was read, and this page's estimates for the newest models sit off those trends: above on the ability lens, below on the scale lens. So the explorer offers a second starting point, the estimated frontier, with the same paces. The levels further up the page give both sets of dates.

Your work on the lineShow

The tasks of every US occupation, and how often each is done, come from , the Labor Department's description of about a thousand jobs, used under its Creative Commons licence. What AI is , for a whole occupation and for each task, comes from , which publishes it against the same list. Both are measured.

How long a task takes is the number the whole page runs on, and nobody publishes it, so it is estimated. Three models from three companies read all 17,578 distinct task statements and each gave a simple, a typical and a demanding length under one written definition. The page keeps the middle of the three middles and the widest low and high, so a task they disagree about carries a wide band. The estimator was checked against , OpenAI's set of 220 real pieces of professional work where the professional recorded how long it really took: on those it ran about a fifth long, and every length here is corrected by that one ratio.

A second pass asks whether a task could be done through a computer at all. Only those are counted against the line, because a thirty-minute task that needs hands is short and no model can do it. That judgment is a model's, made the same way, and it is the weakest link in the section.

Measured, derived, estimated, judgedShow

Measured means a number published by the organization that produced it: 's compute figures, which are themselves estimates with a stated confidence, 's , the scores, results. Drawn solid.

Derived means arithmetic I did on published numbers: the , the trend fits, the lines, the dates a scenario reaches a level, the electricity sum, and the expected score in the chess gap.

Estimated means a figure for something nobody has measured, worked out from published data with a stated method and range: GPT-6 Astra's , the task horizons of models METR has not measured, the marked on the index lens, and the profile's rings. .

Judged means my own call: the starting year of a skill, which events count as stories, whether a task could be done through a computer, and everything the page says about what a level of ability would mean, on the and beyond it. The guesses about the economy, science and society are drawn in violet with a dashed edge, and each names the published forecasts it can be compared with.

What this page cannot tell youShow

Whether any of these systems understands anything. When a machine will match a person at everything, or whether that is the right question. Whether computing will keep turning into ability at the rate it has.

It can tell you what was measured, by whom, how fast it has moved, and what would have to be true for it to keep moving. For deciding what to hand to a machine this year, the reliability gap and the memory gap matter more than any date on the chart.

The data

The files are the measured series exactly as this page uses them. If you see an error in the curated layer, the or the skills, I would like to know.

Update log

Sep 20, 2026

A long page needs to say what each part is for. Every section now opens the same way: its number and how many there are, a line on what it shows and, where there is something to do, a line on how to use it. Every section closes with the point in one line, beside the way onward. A card at the top lists all of them, and the whole list lives in one file, so a section cannot drift from its own description. Five things that were words are now pictures: the profile as a wheel with a person as the outer ring, the creatures on a scale of computing with the year a passed each one, the measured drawn on a log scale and a plain one side by side, the top ten as a leaderboard, and five clock faces you can read yourself. The tables those pictures replaced are still here, folded underneath. Where the opening line of a section said what the section header already said, it is gone, so the full version is shorter than it was this morning despite everything added to it.

Earlier updates (4) ▾
  • Sep 19, 2026

    Your work on the line: pick any of about a thousand US occupations and see the tasks it is actually made of, placed on the same axis as the rest of the page, with the line drawn where the machines are and a year you can drag. The tasks, and how often each is done, come from , the Labor Department's description of work. What AI is , by occupation and by task, comes from and is measured, not projected. How long each task takes is the number nobody publishes, so it is estimated: three models from three companies read all 17,578 distinct task statements, the page keeps the middle of the three, and the estimator is checked against , the one public set of work where professionals recorded the real time. A second pass asks whether a task could be done through a computer at all, because a short task that needs hands is not inside any line. The guesses in What it would mean were rewritten with sharper numbers and new sources the same day.

  • Sep 19, 2026

    Later the same day. The chart now sits where both versions share it, the short version's stories open on it in place, and a story's own address lands on its story. A new story covers OpenAI's report of September 8 of a -checked proof that the can break down, found by running a model it has not released; the Clay Institute still lists the problem as active. What it would mean now reaches past the : for each level of work, for the economy, for science and for society, marked as my judgment and set beside published forecasts with their dates, under a chart of when each level could arrive. The page has a look of its own, with numbered sections, a ruler and warm paper in light mode, and the full version is shorter. The measured series were read again that afternoon and had not changed.

  • Sep 19, 2026

    A short version, about five minutes, now opens for a first visit, with a toggle to this full version. The estimate of for models has not measured now uses a line fitted on the reasoning era only, the fifteen models released since September 2024 that have both numbers. It matches 's own rule of thumb, about five for each doubling, and it puts METR's one attempt at a summer model inside its range. GPT-6 Astra's estimate moved from about 5 to about 4, likely 2.5 to 6, so a work week now reads as in the estimate's range. One rule holds for every level: reached by estimate only when the low end of the range clears it. The section is a table of the top ten models, dates replace words such as today and this month, and every term on the page opens a note, checked by a script. The three measured series and every hand-typed figure were read again on September 19, 2026, and none of the measured numbers had changed.

  • Sep 18, 2026

    First version, on three measured series: 's through September 9, 2026, 's , whose latest clean measurement is of a model from April 2026, and Epoch's through GPT-6 Astra on September 3. The index is also a lens of its own in the explorer. Models METR has not reached carry an estimated horizon , and GPT-6 Astra carries a compute estimate worked out from its disclosed hardware. The chess run to the top of the list and 's September list, both read on September 18. Every dated claim was checked against its live source the same day.

I refresh the three measured series every quarter, and check every dated claim against its live source. The trends, the lines and every date on the page are recomputed from the data when the site is built.

In one line

Every figure here is measured, derived, estimated or judged, and the page says which one it is wherever it appears.

Plain words for every term

Terms and concepts

What this shows

All 129 terms on the page in plain words, grouped, with a box to find one.

Every term with a dotted underline on this page opens a short explanation where it stands. Open a group, or type a few letters to find one.

129 terms, in 13 groups. Open a group to read its terms.

In one line

No word on this page is meant to need a background in the subject. If one did, it is explained here.

If you are working out what to hand to a machine in your operation this year, and what to keep with people, I would like to hear how you are deciding.