The fifteen problems · 05 of 15 · encoding
Byte slicing mojibaked the native-speaker review bundle
The review bundle computed surrounding text as roughly 400 bytes either side of the cited span. Ethiopic is three bytes per code point, so the slice cut mid-character. 396 of 663 context strings, 59.7 percent, contain replacement characters, visible in the very first record.
The review bundle for a native speaker held 663 statements from 43 Amharic articles:
| Relation type | Statements |
|---|---|
| statement | 204 |
| transaction | 96 |
| regulation | 92 |
| quantity | 87 |
| other | 62 |
| temporal | 45 |
| employment | 36 |
| location | 27 |
| ownership | 14 |
It computed surrounding_text as roughly 400 bytes either side of the span. Ethiopic is three bytes per code point, so the slice cut mid-character: 396 of 663 context strings (59.7 percent) contain U+FFFD replacement characters, visible in the very first record. The extractions themselves are clean; the display context handed to the human reviewer is not.
In the same bundle 153 of 663 cited passages are shorter than the statement drawn from them (the English corpus rate was 5 percent), and the review instruction had to separate fidelity from truth explicitly with a five-way scale: says the same, changes the meaning, not in the passage, wrong passage, cannot tell.
× What it broke
The one human review the programme produced was handed context it could not fully read. The bundle carries the mojibake and the short-passage defect together.
✓ What it established
Slice by code point, never by byte, in any script outside ASCII; and separate "is this faithful to the passage" from "is this true" before asking a reviewer anything.
Who reviewed: a native Amharic speaker who is also a data scientist. The bundle is with the reviewer, who is working on the validation as of 22 August 2026; no reading has been returned yet. All Amharic extraction is frozen behind it under the standing rule R6, "do not re-extract the Amharic documents", because re-extracting destroys the thing being reviewed.
user-beadwork/onboarding/amharic-extractions/ · u--ra9.1 · about 2026-08-09 (bundle assembled) · retrieved 2026-08-21
Ask your AI about this page
Paste this page's link into ChatGPT, Claude, or any AI assistant and ask your question in your own words. Every page here publishes a machine-readable copy, so your assistant can read the record directly:
https://ethiopia-build.stoagen.com/problems/05/
Every page here has a markdown twin; this page's is https://ethiopia-build.stoagen.com/problems/05/index.md (also served with .txt appended). The whole site is mapped in one small file at https://ethiopia-build.stoagen.com/site_guide.txt, https://ethiopia-build.stoagen.com/llms.txt describes how the record is organized, and https://ethiopia-build.stoagen.com/agents/ carries the site's notes for assistants.
Published . Last updated . Times come from this page's revision history and can be checked against it.