AGI without the fog: five signals worth watching

This is a live Practicum page for Volume 2, Chapter 10, "The Future in 2030: Four Scenarios." Verified: 2026-08-18. Updated: 2026-08-18.

AGI is easy to discuss in vague terms: when general intelligence will arrive, whether it will replace everyone, or whether a singularity is coming. That does little to help with a decision at work.

This page takes a different approach. Instead of watching the AGI label, track five signals you can test in your own industry, company, or profession.

1. Autonomous task horizon

Question: How long a task can AI complete without constant supervision?

What to watch How to test it in your work
15 minutes short draft, search, classification, revision, or email preparation
1 hour connected analysis, document review, prototype, or meeting plan
Half a day several linked steps: find data, make decisions, produce an artifact, and check it
A day or more difficult work involving context, people, consequences, and unclear criteria

Do not confuse task horizon with the time a model spends running. Ask how much human work, with a clear expected result, the system can finish at an acceptable level of quality.

2. Autonomy

Question: What can AI do by itself, and where must it stop?

Level Example Rule
Adviser suggests options a person chooses
Draft producer prepares material a person reviews and signs off
Operator acts inside a process tests, limits, and a log are in place
Agent chooses its own steps and tools permissions, an owner, and a stop rule are required

As autonomy rises, a good prompt is no longer enough. You need access rules, an action log, tests, and a person who can stop the process.

3. Uneven capabilities

Question: Where does the system look strong but fail without warning?

AI can perform extremely well on one class of tasks and poorly on a nearby one. Do not assume that a successful demo will transfer to your process.

Test three areas:

Area What to include in the test
Normal cases work that happens every day
Edge cases incomplete data, contradictions, poor input, and unusual customers
Red zone cases where AI must stop instead of guessing

A test made only of attractive examples is a presentation, not a test.

The 2026 example worth carrying: ARC-AGI-3

The clearest public demonstration of uneven capability arrived this year. ARC-AGI-3 is an interactive reasoning benchmark: an agent has to explore an unfamiliar environment, work out the rules, and act. As of March 2026, people solve 100% of its environments. Frontier AI systems score below 1%.

Hold that next to the same systems producing expert-level work on roughly half of real professional tasks in GDPval. Both measurements are correct. They are measuring different things: applying learned competence to a familiar kind of task, versus working out the rules of something genuinely new.

There is an honest caveat in the other direction. The predecessor, ARC-AGI-1, also resisted models for years, through a 50,000-fold scale-up, and then moved sharply once test-time adaptation methods appeared in late 2024. So the low score is evidence about today, not a permanent ceiling. The people who build this benchmark deliberately design a new one each time models catch up, which means "AI cannot do ARC-AGI" will keep being true and keep meaning something different.

How to use it: when a vendor says a system reasons, ask which of the two things they measured.

4. Evaluation maturity

Question: Do you have a small evaluation set of your own?

A minimum evaluation for one process includes:

  1. 20 to 30 real tasks.
  2. The expected result or quality criteria.
  3. A review of errors.
  4. The cost of an error.
  5. A decision: can AI act alone, does a person need to review it, or should the task remain manual for now?

Strong general benchmarks do not settle this question. Test the model on your data, your edge cases, and the consequences you will have to own.

The constraint nobody put on the roadmap

One more signal belongs here because it is physical rather than algorithmic. Through 2026 the supply of electric power for AI data centers overtook the supply of AI accelerators as the limiting factor on how fast frontier systems can be trained. Gartner puts global data-center electricity demand above 1,000 TWh in 2026, about double the 2023 baseline, and transformer lead times now run to roughly two and a half years. See who owns the rails.

Whatever you believe about AGI timelines, they now run through a grid connection queue. A forecast that does not mention power is a forecast about software written by someone who has not checked what it plugs into.

5. Readiness of people and institutions

Question: Does AI make people more independent, or more passive?

Signal Good sign Bad sign
School and training students learn to check, explain, and challenge AI AI is either banned or used to hand out finished answers
Work an employee builds a workflow and understands its failure modes an employee copies output without checking it
Hiring candidates may use AI and must explain their reasoning the test only checks work done without the tool
Company rules define which tasks AI may perform the company has subscriptions but no clear responsibility

Powerful AI does not make an organization mature. It reveals who can learn and verify, and who is waiting for instructions.

Quick quarterly review

Once a quarter, choose one process and complete this table.

Process Task horizon Autonomy Where it fails Evaluation set? What will change?
15 min / 1 hour / half day / day+ adviser / draft producer / operator / agent yes / no

If the evaluation column says "no," do not raise the level of autonomy. Build the test first.

Sources

Version: 2026-08-18. Next review: before the final publication of Volume 2 or after a major update to frontier model evaluations.

AGI without the fog: five signals worth watching