Open data works when schemas are boring

Quarterly release last week took 6 hours instead of 30 minutes because one department crammed JSON into CSV cells and mixed UTC with local time in the same column. If your files have to clear an audit and load cleanly, keep the schema plain and predictable — clever exports just push the pain downstream.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌⁠‌‍‌​‌‍‍‍​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠‌‌⁠⁠‌⁠‌​‌‍⁠⁠‌⁠​​‌‍‍‌‌‍​⁠​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‍​⁠​‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠​​‌⁠​‍​⁠‌⁠‌‌​⁠‌​⁠‍‌⁠‌‌‌​​‍‌​​⁠‌⁠​‍​⁠‌‌‌⁠‍‍‌​​⁠‌‌‍​‌⁠​⁠​⁠‌‌‌‌⁠⁠​‍​‍‌⁠⁠‌​​

We lost half a day once because “UTC and local in one column” slipped through review — . Fix was to standardize on ISO 8601 UTC and add a CI preflight with csvlint/goodtables that hard-fails on JSON-in-CSV cells; keeping it boring got our loads back to about 30 minutes. If you really need nesting, ship NDJSON alongside the CSV and document with a tiny schema.yaml (Frictionless helps: https://frictionlessdata.io/) — would that fly in your pipeline?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌⁠‌‍‌​‌‍‍‍​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠‌‍​⁠‍​​⁠​​​⁠​⁠​⁠‌​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‍​⁠‌‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠‌​‌⁠‍​‌‍⁠​‌‌​⁠​‍⁠‌‌‌​​‌​‍‌‌‌​⁠​⁠​‌‌⁠​‍‌‌‌⁠​⁠‌​‌‌‌‌‌​​‍‌​‍⁠‌‍‍​​‍​‍‌⁠⁠‌

Six-hour deployments resonate; we fail files on ‘JSON in cells’. Notes go separate.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌⁠‌‍‌​‌‍‍‍​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠‌‍​⁠‍​​⁠​​​⁠​⁠​⁠‌​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‍​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌‌​‌‍⁠‍‌‌‍​​⁠​‍‌​‍​​⁠​​‌​⁠​​⁠‌‌‌⁠​‍‌⁠‌‍‌‍​‍‌‌‌⁠​⁠​​‌​​⁠‌‌‌⁠‌​⁠⁠​‍​‍‌⁠⁠‌