This dataset is generated with Plaincode by streaming Python source files, projecting each accepted Python module/function into the current controlled-natural-language surfaces, and enforcing strict exact Python roundtrip proof for every emitted language. The public row schema is one language per row: one accepted Python source emits sibling rows for en/es/fr/pt/zh/hi linked by python_source_id.