CONTEXT BEFORE CONCLUSIONS

How to read the results.

This is an archive of model-generated artifacts, with the experiment setup kept visible.

What is being tested?

Models are asked to produce single-file HTML games, tools, and visualizations from task prompts. A small number of tasks produce other source formats. The archive lets you inspect and play their outputs and explore how the available authoring tools affect what they produce.

The authoring tiers

Basic
A single response containing the complete HTML. No authoring or verification tools.
Middle
Artifact-scoped write, edit, and read tools, with no lint, outline, snapshot, or browser feedback.
Full
Artifact editing plus browser verification, behavioral probes, navigation, and checkpoints.
Composer
Structured JavaScript units, CSS blocks, and DOM edits assembled into one HTML file, alongside verification tools.
TypeScript
Typed ES modules compiled and linked into the artifact. Compiler diagnostics provide feedback; verification still checks the browser behavior.
Legacy
An older result whose authoring tier was not recorded in the current folder structure.

These descriptions reflect the current personas. Historical configurations may differ. Prospective runs retain persona hashes and package versions.

Prompt revisions and attempts

The current prompt on a task page is a reference. Historical artifacts often lack a snapshot of the exact instructions used. These results show an unknown prompt revision rather than being attributed to today's text.

New runs capture task text, persona source hash, package versions, and whether an existing artifact was provided as input. A continuation has a different meaning from a first attempt. Effective runtime system instructions are not currently captured, and additional private instructions are labeled when present.

Reviews

Human judgements use separate 1–5 dimensions. Null values mean unscored. There is no composite quality score. Broken outputs remain available in the archive; their broken flag is a reviewer judgement, separate from browser errors.

Visual design
Aesthetics, polish, layout, typography, color. At 1: Ugly or broken layout. At 5: Polished and cohesive; could ship.
UX / feel
Controls, feedback, responsiveness, discoverability. At 1: Frustrating or confusing to use. At 5: Responsive, intuitive, great feedback.
Works
Does it function correctly when actually exercised. At 1: Doesn't run or core loop broken. At 5: Everything functions as intended.
Prompt spirit
Captures what the prompt was going for, with taste. At 1: Built the wrong thing. At 5: Nails the intent with creative taste.
Delight
Wow factor — would you show this to someone?. At 1: Nothing memorable. At 5: Would show someone.

Agent judgements are shown separately by judge model, rubric version, and judge run. Historical reviews tied only to file paths have an unverified association with exact artifact bytes. New hash-associated reviews avoid carrying a score over when a working file is overwritten.

Costs and generation metrics

Token counts, reported cost, turns, and tool failures come from the captured metadata. Unknown values stay unknown. For continued work, these metrics may cover only the last session rather than all effort that contributed to the artifact. Current prices are not used to estimate old costs.

Playing the artifacts

Each HTML result runs in a separate, sandboxed frame. External scripts, module imports, styles, fonts, assets, and fetches can load from any HTTPS host, including future library providers. Plain HTTP dependencies and embedded frames remain blocked. CDN availability, CORS, and obsolete library URLs can still affect playback. Prompt constraints such as “no libraries” are evaluated as part of the result, rather than enforced by changing playback. Persistent storage is enabled on the separate artifact origin, which cannot access gallery storage. Artifacts on that origin share save storage. Clipboard access and other browser features may be unavailable.

The downloadable original preserves the generated bytes. The screenshot is an initial-state browser capture, not proof that every interaction works. Random or animated outputs can look different on another run. Browser diagnostics are technical observations, not scores.

What comparisons can tell you

Compare exact prompt revisions where possible and check authoring tiers, additional instructions, and attempt histories. Different providers, settings, and sample sizes can affect results. This collection supports exploration; it does not establish a definitive ranking of models.

Publication and history

Generation dates and publication dates are shown separately. The archive retains immutable versions even if local working files later change. Scores can be updated without creating a new generation result. Selected results may be withdrawn; their old result links show an archive notice.