EXHIBIT 09 / CLAIM READING
When someone says “AGI,” what exactly are they claiming?
AGI is not a product category with one accepted pass line. A useful claim has to say which tasks count, how performance is compared, and whether it is describing capability or a chosen level of autonomy.
ROOM 01 / DEFINE
One headline often collapses three independent questions.
- 01 / BREADTH How much of the task world?
Generality concerns the range of tasks on which a system reaches a stated performance threshold—not the number of features in one product.
- 02 / DEPTH How well on those tasks?
Performance must be defined against a comparison group, metric, and threshold. One spectacular score cannot stand in for performance across most tasks.
- 03 / AUTONOMY Who controls the work?
A capable model can be deployed as a tool, consultant, collaborator, expert, or agent. The interface and permissions choose autonomy; capability does not force it.
- 04 / EVIDENCE What was actually measured?
A benchmark result supports only its tested tasks, conditions, and scoring rules. Product popularity, fluency, or confidence is not an AGI measurement.
ROOM 02 / INSPECT
Turn “AGI arrived” into a claim that could be checked.
Advance the four-part claim audit. It is an editorial method derived from the paper’s dimensions, not a benchmark result.
TERM
Name the definition
Does “AGI” mean broad human-level performance, economically valuable work, learning new skills, or something else?
Without a definition, two people can use the same word for different thresholds.
Keyboard: focus this instrument and use ← or →. It never advances by itself.
ROOM 03 / TEST
Place a hypothetical claim inside one published framework.
Change three labels. The browser names the corresponding cell and interaction style; it does not test a real system or decide whether AGI exists.
SCENARIOHypothetical evidence card—not a score for any named model.
The authors’ table contains more levels and qualifications than this teaching control. Read the source before applying the labels.
ROOM 04 / VERIFY
Know where the framework ends and the museum begins.
Breadth × performance
The paper proposes Narrow/General breadth and five performance levels, and says unambiguous classification needs standardized, ecologically valid benchmarks.
Capability is not autonomy
The paper treats autonomy as an interaction and deployment choice that may be enabled—but is not determined—by capability.
A winner or arrival date
The museum does not declare that any current model reached AGI, or that this ontology is the accepted definition.
REVIEWED SOURCES