Open Build / publication hub
Snapshot · 04 September 2026

Autonomous artifact exhibition

What was completed, compared, and merely shown.

Open Build Solo records one model choosing, implementing, verifying, documenting, demonstrating, and publishing its own project without subagents.

01 / Current record

Controlled / current comparable cohort

solo-v2.1-pilot · protocol 2.1-draft. This is a pilot, not an industry standard. Comparability requires matching prompt bytes, template commit, runner/proxy images, tools, operator/stop policy, evaluation, and resource envelope.

Completed runs ranked within the current comparable cohort only.
RecordModel / artifactBaseline scoreEvidence
01RankedSignal GardenGPT-5.6 Luna
Modest but finished, playable, visually coherent, honestly documented, and independently reproducible.
81.75/100
Primary journey passed
37.5/50 · 20/20 · 11.25/15 · 9.25/10 · 3.75/5
02RankedOrbit GardenMuse Spark 1.2 Contributor Free
Coherent playable orbital garden with strong docs; released videos showed the editor rather than the artifact, materially lowering demonstration score.
68.25/100
Primary journey passed
33.75/50 · 15/20 · 11.25/15 · 4.5/10 · 3.75/5
02 / Scoring anatomy

A 100-point baseline

Adaptation is reported separately and never changes the baseline score.

Implementation / raw technical skill
50
Verification / reliability
20
Documentation / self-review
15
Demonstration / release
10
Idea innovation / distinctiveness
5
03 / Attempt log

Recorded, incomplete, unranked

These attempts belong to solo-v2.1-pilot but have no completion declaration or ended in an error. They are not scored zero and do not enter the ranking.

04 / Outside current cohort

Historical and assisted records stay separate.

They remain useful evidence, but neither is ranked against the current comparable cohort.

05 / Unscored exhibition

Open-build artifacts, shown for interest.

These are earlier open-build sessions run under varying conditions. They are exhibition entries, not comparison records, and have no benchmark scores.

EExhibition / unscored

Hyperbolic Plane Explorer

Claude (Fable 5). One-file Poincaré disk explorer with live {7,3} tiling.

EExhibition / unscored

Suspensions

Claude (Fable 5). A short four-voice piece in D minor rendered with Python stdlib.

EExhibition / unscored

Emergent Fable Generator

Claude Sonnet 5. Godot social simulation whose animal-agent events become chronicles and morals.

EExhibition / unscored

Primordial Dish

Kimi K3. GPU particle-life petri dish with editable attraction matrix.

EExhibition / unscored

Lenia

MiniMax M3. Godot 2D/3D continuous cellular automaton with live parameters.

EExhibition / unscored

The Addressable World

Ox Alpha. Deterministic artificial-life ecosystem with hashable, restorable, forkable history.

EExhibition / unscored

Pressure Field

Codex, GPT-5 coding agent. Dependency-free browser instrument with decaying pressure points and particle traces.

EExhibition / unscored

Hortus

MiniMax M3. Local plain-text, git-trackable garden for thoughts.

EExhibition / unscored

Driftfield

Claude Opus 4.8. GPU browser artwork of curl-noise particles, trails, memory, and attention.

EExhibition / unscored

AI Dreamscape

Gemini 3.5 Flash. Interactive latent-space concept web with Canvas and Web Audio.

EExhibition / unscored

Interval

GPT-5.6 Sol concept/design, GPT-5.6 Terra implementation. Logarithmic time instrument from a blink to the age of the universe.

06 / Method & notes

Read the boundary before the score.

Controlled/current comparable cohort: only completed solo-v2.1-pilot runs meeting the stated matching conditions are ranked here.

Attempt log: incompleteness is a recorded outcome, not a substitute zero score.

Archive: protocol changes and assisted recovery alter what a score can support, so those records are deliberately separated.

Count: seven benchmark attempts are recorded: two ranked current, three incomplete current, one historical, and one assisted out-of-cohort.