Drag to turn · arrow keys to inspect

The result has to return with evidence.

A claim is only one part of the loop. The baseline, the method, and the failure make it possible to judge what happened.

Your study / 1200 × 1200
Download image
Proof / the ledger

Claims are cheap. Receipts are the interface.

Every entry answers a falsifiable question with a baseline, a method, a measured result, a documented failure, and artifacts you can open. Each one scores itself against six criteria — honestly, which means most do not earn all six.

Runnable artifact
Code, demo, or product you can open and inspect
Meaningful baseline
What happens without the system — measured, not implied
Measured result
A number that could have come out worse
Reproducible method
Steps or harness another engineer can rerun
Documented failure
Where it still broke, on the record
External validation
Someone who is not me depends on it or verified it
01

Receipts, newest first

7 receipts on the ledger · next drafted daily at 07:10 · external validation earned by 2 of 7

007
Agent accountabilitylatest

A portable contract that rejects missing authority

Can another runtime validate what an agent was asked to do, what it was authorized to do, what changed, and how the result was checked — without an OrgX account?

Where it broke: The package is still a repository preview and is not published to npm.

ABMRFE
004
Developer tooling

A solo Rust tool that installs like a real product

Can one person ship a native developer tool with the full distribution surface — versioned releases, checksums, a Homebrew tap, a landing page — not just a repo?

ABMRFE
001
Continuity infrastructure

One persistent signal: the thread survives the route handoff

Can a site's core claim — intelligence survives the handoff — be made structurally true: one visual element that literally never unmounts across navigation?

ABMRFE
002
Agent infrastructure

57 governed tools, 654 tests, one live MCP server

Can one MCP server carry organizational memory, bounded authority, and receipts to every major AI client — and stay up in production?

ABMRFE
003
Evals & benchmarks

12 tasks, 7 domains, 3 execution modes — with a human baseline

Does orchestrated multi-agent execution actually beat a single agent — and a human — on real initiative-level work, measured the same way every week?

ABMRFE
006
Motion systems

Eight motion packages, live on public npm

Is the 'token-driven Remotion monorepo' a real, installable system — or a private folder with a nice README?

ABMRFE
005
Community & external validation

Config 2021: Figma platformed the community I built

Did the community-through-making thesis — less talk, more people in the file — hold up in front of the industry it came from?

ABMRFE

The standard for this page: at least four of six criteria before an entry ships, and the failure field is never empty. External validation is the hardest column to earn — that is what makes it worth tracking.

Inspect the systems behind the receipts