A public engineering record · 14 July to 12 August 2026

Building the Ethiopia knowledge base

Between 14 July and 12 August 2026 a two-person team, one giving the agents their orders and a native Amharic speaker who is also a data scientist, working with a team of AI agents, tried to build a machine-readable knowledge base about Ethiopia from public sources and to extract knowledge graphs from it. This site is the record of that attempt: where the data came from, how the graphs were built, and what went wrong, with the Amharic problems given the prominence they earned.

1,153

Wikipedia captures, hashed and revision-stamped

1,250

World Bank documents acquired with full provenance

15 of 15

graphs different: three documents, five model draws each, not one document reproduced its own graph once

On 11 August 2026 the operator ruled the Ethiopia corpus "nowhere near worked", pulled the Ethiopia programme out of the architecture, and made the architecture generic so corpus work could start over. This site exists because the record of why is worth more than the corpus was.

user-beadwork/plans/PLAN_lens-graph-similarity_2026-08-11.md · v1 to v3.2 · retrieved 2026-08-21

The record, in six chapters

01 Where the data came from

Five source lanes, fifty audited families, and the table of hosts that blocked, timed out, or lied about their encoding.

02 How the data was ingested

Immutable captures, monotone ingestion, the trusted execution broker, and why Obsidian was demoted.

03 Three generations of graphs

Deterministic claims, OO-LD hypergraphs, the lens-graph architecture: what each produced and why each was replaced.

04 The fifteen Amharic problems

The centrepiece. Concrete failures specific to a low-resource language with its own script, in the order a builder meets them.

05 Timeline and open items

14 July to 11 August, dated from tickets and commits, with an honest state column for everything still open.

06 Who did the work

Fifteen agent seats by role and runtime, what each produced, and what they recorded against themselves.

ethiopia-program/corpora/ethiopia/ · u--2x9 · eth-5q7 · retrieved 2026-08-21

commit 3fdb48ab · the convergence artifact, deleted with the corpus, survives only in git history

How to read this site

Every figure on these pages carries a repository path, ticket id or commit in an evidence strip like the two above, so it can be re-verified against the archive. Failures are content here: blocked sources, rejected runs and measurements that came back zero are listed with the same care as successes. People appear by role only. Ethiopic text renders from your device's system fonts and is marked reviewed or awaiting review. Nothing on this site is triumphant and nothing apologizes; it is a lab notebook, published.

Two further pages sit outside the six chapters: the evidence map, which lists every repository, document and ticket store the record draws on, and agent access, the site's notes for the AI assistants that read it.

Ask your AI about this page

Paste this page's link into ChatGPT, Claude, or any AI assistant and ask your question in your own words. Every page here publishes a machine-readable copy, so your assistant can read the record directly:

https://ethiopia-build.stoagen.com/

Published . Last updated . Times come from this page's revision history and can be checked against it.