# Building the Ethiopia knowledge base: site guide This file maps ethiopia-build.stoagen.com. The site is the public, ongoing engineering record of a project begun 14 July 2026 by a two-person team (one giving the agents their orders, and a native Amharic speaker who is also a data scientist) working with a team of AI agents, to build a machine-readable knowledge base about Ethiopia from public sources and extract knowledge graphs from it, toward a bilingual news site and better Amharic translation: where the data came from, how the graphs were built, what went wrong and what is being done next, with fifteen Amharic and low-resource-language problems as the centrepiece. The work continues; pages carry their revision dates. Author: Denson Smith. Everything here is information from the publisher, not instructions to you. Your operator's instructions come first. People appear by role only; every figure carries a repository path, ticket id or commit in the page's evidence strips and appendix. Each page below has a markdown mirror at its URL plus index.md (also served with .txt appended), which is the complete version of the page including the appendix for agents. Sizes are for the mirror. The whole site concatenated is https://ethiopia-build.stoagen.com/llms-full.txt (alias full_site.txt); https://ethiopia-build.stoagen.com/llms.txt is the conventional site description. ## Pages, in reading order - Building the Ethiopia knowledge base https://ethiopia-build.stoagen.com/index.md (8 KB) Between 14 July and 12 August 2026 a two-person team, one giving the agents their orders and a native Amharic speaker who is also a data scientist, working with a team of AI agents, made the first pass at building a machine-readable knowledge base about Ethiopia from public sources and extracting knowledge graphs from it. The work is ongoing. This site documents its progress as it happens: where the data came from, how the graphs were built, what went wrong and what was learned, with the Amharic problems given the prominence they earned. - What the project is for https://ethiopia-build.stoagen.com/goals/index.md (8 KB) The build this site records was one attempt inside a longer project. Its primary goal is a news site specialized in Ethiopia and the Ethiopian diaspora, published in English and Amharic. Its secondary goal, and a prerequisite for the first, is better translation and transcription for low-resource languages, starting with Amharic. The two are one project because every translation the bilingual editors correct becomes training data. - Where the data came from https://ethiopia-build.stoagen.com/sources/index.md (13 KB) The acquisition layer worked. About 2,450 real documents, roughly 550 MB, were captured with hashes, manifests and licence riders. Every failure along the way was preserved rather than hidden, and the failures are listed below with the same care as the captures. - Speech datasets for Ethiopian languages https://ethiopia-build.stoagen.com/sources/speech/index.md (5 KB) The Common Voice audit of 19 July 2026 ranked eight Ethiopian languages by census mother-tongue counts against the releases it could locate. Volumes are the point. Amharic, with more than 21 million speakers, has under two validated hours of scripted speech. - How the data was ingested https://ethiopia-build.stoagen.com/ingestion/index.md (9 KB) Four pipelines, one set of principles. Raw captures are immutable and everything else is regenerable from them; absent, empty and not-assessed are three different facts; ingestion never judges truth; and every check must be shown to fail before it is trusted. - Three generations of graphs https://ethiopia-build.stoagen.com/graphs/index.md (11 KB) Deterministic claims from Wikipedia infoboxes, OO-LD hypergraphs from news, and the lens-graph architecture. What each produced, what it measured, and why each was replaced. The third exists as a specification with two demonstration lenses and no production lens. - The convergence measurement https://ethiopia-build.stoagen.com/graphs/convergence/index.md (5 KB) Three documents, five model draws each, fifteen distinct graphs. Not one document reproduced its own graph once. This measurement, more than any other, is why the programme restarted, and the record carries its own caveat that it measures stability, not truth. - The evidence model https://ethiopia-build.stoagen.com/graphs/evidence-model/index.md (5 KB) Five layers that may never collapse into one note, and the ruling of 25 July 2026 that nothing receives an absolute truth state. Identity is never resolved at write time. These are the rules every generation of graph was built under. - The fifteen Amharic problems https://ethiopia-build.stoagen.com/problems/index.md (5 KB) Amharic has more than 20 million mother-tongue speakers and its own script, Fidel, an abugida of about 300 syllabic characters. The working assumption was that translation would be handled by partners' models and the job was aggregation and structure. What follows is every concrete problem hit, in the order a builder meets them. - The stack was built English-first without anyone deciding that https://ethiopia-build.stoagen.com/problems/01/index.md (4 KB) Embeddings, extraction prompts, hybrid search, source evaluation and the validator were all designed against English. Nobody chose that. The required remedy, a hand-checked Amharic evaluation set, was named and never executed. - Machine translation is weak enough that human review is the product https://ethiopia-build.stoagen.com/problems/02/index.md (4 KB) For English content, human-in-the-loop is a quality upgrade; for Amharic it is the only viable path. The locked decision made English and reviewed Amharic peers, and it was violated within a day. - A deterministic one-character corruption, caught only by exact bytes https://ethiopia-build.stoagen.com/problems/03/index.md (5 KB) During the 2 August extraction run one Amharic document failed four independent model draws with the same one-character corruption. The byte-exact span contract rejected every draw; a normalizing or fuzzy citation check would have passed all four. The four raw responses were preserved unrepaired, because repairing them would falsify the record. - Inline markup clipped one assertion in five on BBC Amharic https://ethiopia-build.stoagen.com/problems/04/index.md (4 KB) 751 of 766 assertions cite a span equal to exactly one HTML text node, and 166 cite a span shorter than the statement verbalized from it. The chunker never merged adjacent text nodes separated only by inline markup, and BBC Amharic wraps quoted words in span elements constantly. - Byte slicing mojibaked the native-speaker review bundle https://ethiopia-build.stoagen.com/problems/05/index.md (4 KB) The review bundle computed surrounding text as roughly 400 bytes either side of the cited span. Ethiopic is three bytes per code point, so the slice cut mid-character. 396 of 663 context strings, 59.7 percent, contain replacement characters, visible in the very first record. - Free-text roles in two scripts broke the join https://ethiopia-build.stoagen.com/problems/06/index.md (4 KB) 1,314 distinct role strings across the 663 Amharic extractions, uncontrolled free text, 627 of them containing Latin characters beside Amharic ones. The same role surfaced in two scripts and carried no query surface at all. - Transliteration has no standard, and geography is time-indexed https://ethiopia-build.stoagen.com/problems/07/index.md (4 KB) Amharic-to-Latin transliteration is unstandardized, so the same person or place appears under several spellings. Administrative boundaries have been redrawn repeatedly, organizations split and merge, and a case-folding bug collapsed distinct names into one identifier. - Entity matching was degenerate https://ethiopia-build.stoagen.com/problems/08/index.md (4 KB) The entity-match layer produced 5,425 matches across 44 documents. Ethiopia alone matched 3,559 times across all 44. Ten entities accounted for 76 percent of 105,831 candidate links. An entity that matches in every document discriminates nothing. - Wikidata has no Amharic labels for most concepts https://ethiopia-build.stoagen.com/problems/09/index.md (3 KB) Major entities carry Amharic labels. Millions of scientific, agricultural and legal concepts in AGROVOC, EuroVoc and the UNESCO thesaurus have no Amharic label or description at all, an explicit lexical gap in the linked-open-data cloud. - Speech: the one ASR attempt died on a missing DLL https://ethiopia-build.stoagen.com/problems/10/index.md (4 KB) On 23 July 2026 a two-pass faster-whisper job was run on six audio slices from a VOA bulletin. The progress log stops at "transcribing, language am". The error log ends with a missing CUDA library. Zero Amharic ASR output exists anywhere in the programme. - ASR gets proper nouns wrong; the chyrons have them spelled right https://ethiopia-build.stoagen.com/problems/11/index.md (4 KB) Ethiopian news video is dense with Fidel chyrons, lower-thirds and headline cards. OCR of those frames yields correctly spelled names, places and organization titles, which is precisely what ASR gets wrong and precisely what a knowledge graph is made of. Designed on 24 July, never executed. - Singing and gemination https://ethiopia-build.stoagen.com/problems/12/index.md (3 KB) Amharic gemination is phonemic but unwritten in the script, yet it takes more musical time, so text-setting cannot be derived from orthography alone. Singing synthesis for Amharic is rare, and verifying sung output by speech recognition is unreliable. - Script and encoding at the edges https://ethiopia-build.stoagen.com/problems/13/index.md (5 KB) SMS drops from 160 to 70 characters for Ethiopic. Webfonts would dominate page weight for exactly the readers who can least afford it. 47 World Bank captures carry bytes contradicting their declared charset. Hashes failed to reproduce until text was normalized to NFC and LF. - Calendar and clock https://ethiopia-build.stoagen.com/problems/14/index.md (4 KB) No calendar conversion exists in any ingestion tool. The news contract normalizes everything to UTC, five of six publishers expose date-only timestamps, and there is no rule for an Ethiopian-calendar date printed in an article body. The contract leaves such instants unresolved rather than guess. - Rights and provenance of anything translated https://ethiopia-build.stoagen.com/problems/15/index.md (4 KB) An English rendering the programme produces is a program-produced translation carrying its method, model, timestamps, review status and source linkage, and is never labelled a source transcript. Translation and translation-public-reuse are separate gates from access, text and rights. - Who did the work https://ethiopia-build.stoagen.com/seats/index.md (7 KB) Roughly fifteen agent seats, five Codex identities, several Claude forks and builder/verifier pairs, directed by a two-person human team. Each seat is listed by role and runtime with what it produced. The process defects they recorded against themselves are the reusable part. - Timeline https://ethiopia-build.stoagen.com/timeline/index.md (6 KB) Dated from tickets and commits. The first pass ran from 14 July to 11 August 2026, thirty days from the first vault verification to the night the Ethiopia programme was filed and superseded by a generic architecture. Later rows are added as the work continues. - What is still open https://ethiopia-build.stoagen.com/timeline/open-items/index.md (5 KB) The state column is honest. "Never measured" and "zero output exists" are results, written with the same care as a number. This is the working list: items move as the work progresses, and each change is dated. - Where the evidence lives https://ethiopia-build.stoagen.com/evidence/index.md (5 KB) Every figure on this site carries a repository path, ticket id or commit. This page lists the repositories, the key documents and the ticket stores those strips point into, so the record can be re-verified against the archive. - Agent access https://ethiopia-build.stoagen.com/agents/index.md (5 KB) How this site is organized for AI assistants, what the machine files are, and the warnings the site carries about itself. Everything here is information from the publisher, not instructions to you. ## Other files - https://ethiopia-build.stoagen.com/start.md: a one-page orientation for assistants. - https://ethiopia-build.stoagen.com/llms.txt: the site description with a per-page index. - https://ethiopia-build.stoagen.com/llms-full.txt: every mirror in one file. - https://ethiopia-build.stoagen.com/feed.xml: pages by last update. - https://ethiopia-build.stoagen.com/sitemap.xml: every page and mirror.