Selecting a source from a bar at the top meant every page was implicitly scoped to a hidden bit of state. The source is now a route parameter: /sources lists them, /sources/:name owns the source, and Import, Rules, Mappings, Records, and Pivot are tabs beneath it. Sources.jsx split into SourceList and SourceDetail, with Section and SampleTable extracted so the create dialog and the detail page share them. The detail page is grouped into titled panels — Connection, Fields and view, Sample rows, Maintenance, Delete — rather than one flat form. Two new top-level pages, following how Monarch separates these: - Import, because importing is frequent and configuring a source is not. Every source in one list, with a sync button for bank feeds. - Bridge, because one SimpleFIN credential covers every account, so connection state belongs in one place rather than per source. Loads only when asked, since it queries SimpleFIN, and totals the balances. The dark mode toggle lived in the status bar and moved to the sidebar. The out-of-sync and reprocess banners are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G2HFeU5neCKagTnmA6o9Tu |
||
|---|---|---|
| api | ||
| database | ||
| docs | ||
| examples | ||
| ui | ||
| .env.example | ||
| .gitignore | ||
| CLAUDE.md | ||
| dataflow.service | ||
| manage.py | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
Dataflow
A simple data transformation tool for importing, cleaning, and standardizing data from various sources.
Point it at a messy CSV — bank transactions, product lists, anything repetitive — and it will deduplicate on import, pull structure out with regex rules, map the extracted values to clean output, and serve the result through a web UI and REST API.
How it works
- Sources define where data comes from and which fields make a record unique
- Rules extract information with regex (
extractorreplacemode) — e.g. pull the merchant out of a transaction description - Mappings turn extracted values into clean output —
"DISCOUNT DRUG MART 32"→{"vendor": "Discount Drug Mart", "category": "Healthcare"} - Records are then queryable, pivotable, and exportable
Each record keeps three layers: data (raw import), transformed (rule and mapping output),
and overrides (manual edits). Reads merge them in that order, so re-running the rules never
clobbers something you typed by hand.
Stack
PostgreSQL with JSONB storage, a Node.js/Express API, and a React SPA served from public/.
HTTP Basic auth, configured in .env.
Getting started
Requires PostgreSQL 12+, Node.js 18+, and Python 3.
npm install
python3 manage.py # interactive setup: .env, database, schema, functions, UI, service
The UI is then at http://localhost:3020 and the API at http://localhost:3020/api
(port set by API_PORT in .env).
For a walkthrough that creates a source, adds rules and mappings, and imports the sample
CSV in examples/, see docs/getting-started.md.
Documentation
| docs/getting-started.md | Tutorial — build a working pipeline from scratch with curl |
| docs/spec.md | Full reference — architecture, schema, data flow, API, manage.py |
| docs/ui.md | Frontend: React + Vite build, key packages |
| docs/perspective.md | Pivot table: pinned versions and API reference |
Project structure
dataflow/
├── manage.py # interactive setup / deploy / uninstall
├── database/ # schema.sql + one .sql file per API route
├── api/ # Express server, routes, auth middleware
├── ui/ # React source (built to public/)
├── public/ # built UI, served as static files
├── docs/
└── examples/ # sample CSV for the tutorial
Both the API routes and the SQL are organized one file per resource, so api/routes/rules.js
and database/rules.sql are the two halves of the same feature.
database/*.sql is the source of truth for every database function — never edit one directly
in the database, or the next redeploy will silently revert it.
License
MIT