# Dataflow A simple data transformation tool for importing, cleaning, and standardizing data from various sources. Point it at a messy CSV — bank transactions, product lists, anything repetitive — and it will deduplicate on import, pull structure out with regex rules, map the extracted values to clean output, and serve the result through a web UI and REST API. ## How it works 1. **Sources** define where data comes from and which fields make a record unique 2. **Rules** extract information with regex (`extract` or `replace` mode) — e.g. pull the merchant out of a transaction description 3. **Mappings** turn extracted values into clean output — `"DISCOUNT DRUG MART 32"` → `{"vendor": "Discount Drug Mart", "category": "Healthcare"}` 4. **Records** are then queryable, pivotable, and exportable Each record keeps three layers: `data` (raw import), `transformed` (rule and mapping output), and `overrides` (manual edits). Reads merge them in that order, so re-running the rules never clobbers something you typed by hand. ## Stack PostgreSQL with JSONB storage, a Node.js/Express API, and a React SPA served from `public/`. HTTP Basic auth, configured in `.env`. ## Getting started Requires PostgreSQL 12+, Node.js 18+, and Python 3. ```bash npm install python3 manage.py # interactive setup: .env, database, schema, functions, UI, service ``` The UI is then at `http://localhost:3020` and the API at `http://localhost:3020/api` (port set by `API_PORT` in `.env`). For a walkthrough that creates a source, adds rules and mappings, and imports the sample CSV in `examples/`, see **[docs/getting-started.md](docs/getting-started.md)**. ## Documentation | | | |---|---| | **[docs/getting-started.md](docs/getting-started.md)** | Tutorial — build a working pipeline from scratch with curl | | **[docs/spec.md](docs/spec.md)** | Full reference — architecture, schema, data flow, API, `manage.py` | | **[docs/ui.md](docs/ui.md)** | Frontend: React + Vite build, key packages | | **[docs/perspective.md](docs/perspective.md)** | Pivot table: pinned versions and API reference | ## Project structure ``` dataflow/ ├── manage.py # interactive setup / deploy / uninstall ├── database/ # schema.sql + one .sql file per API route ├── api/ # Express server, routes, auth middleware ├── ui/ # React source (built to public/) ├── public/ # built UI, served as static files ├── docs/ └── examples/ # sample CSV for the tutorial ``` Both the API routes and the SQL are organized one file per resource, so `api/routes/rules.js` and `database/rules.sql` are the two halves of the same feature. `database/*.sql` is the source of truth for every database function — never edit one directly in the database, or the next redeploy will silently revert it. ## License MIT