Go to file
Paul Trowbridge 3613037ab5 Add SimpleFIN Bridge sync as an alternative to CSV import
Sources with a `simplefin` block in their config can pull transactions
straight from the bridge instead of taking a CSV upload. Only the fetch
differs — dedupe, logging, and transformation reuse the import path.

The access URL is the whole credential, so it lives in .env rather than
the database that manage.py offers to reset. Claiming a setup token is
exposed as an endpoint because the token is single-use and easy to burn.

The bridge answers 200 with a populated `errors` array when a bank is
failing, which would otherwise read as a successful empty pull — those
errors ride along in the sync response and show on the Import page.

Pending transactions are skipped by default: they get a new id once they
post, which would import the same charge twice under two keys. Sources
should use ['id'] as constraint_fields — the transaction id makes
overlapping pulls free while keeping genuinely repeated charges distinct.

Verified against a stubbed bridge response, not a live account.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G2HFeU5neCKagTnmA6o9Tu
2026-08-01 12:47:04 -04:00
api Add SimpleFIN Bridge sync as an alternative to CSV import 2026-08-01 12:47:04 -04:00
database List override and transformed keys as source fields; show record id 2026-07-26 23:10:01 -04:00
docs Add SimpleFIN Bridge sync as an alternative to CSV import 2026-08-01 12:47:04 -04:00
examples Consolidate documentation into docs/ and cut the duplication 2026-07-26 21:55:58 -04:00
ui Add SimpleFIN Bridge sync as an alternative to CSV import 2026-08-01 12:47:04 -04:00
.env.example Add SimpleFIN Bridge sync as an alternative to CSV import 2026-08-01 12:47:04 -04:00
.gitignore Track package-lock.json for both the API and the UI 2026-07-26 21:54:10 -04:00
CLAUDE.md Consolidate documentation into docs/ and cut the duplication 2026-07-26 21:55:58 -04:00
dataflow.service Add unified deploy.sh and systemd service unit 2026-04-05 15:53:02 -04:00
manage.py Flatten database/queries into database/ and fix five stale functions 2026-07-26 21:55:37 -04:00
package-lock.json Track package-lock.json for both the API and the UI 2026-07-26 21:54:10 -04:00
package.json Remove obsolete deploy scripts, migration helpers, and unused UI assets 2026-07-26 13:15:46 -04:00
README.md Consolidate documentation into docs/ and cut the duplication 2026-07-26 21:55:58 -04:00

Dataflow

A simple data transformation tool for importing, cleaning, and standardizing data from various sources.

Point it at a messy CSV — bank transactions, product lists, anything repetitive — and it will deduplicate on import, pull structure out with regex rules, map the extracted values to clean output, and serve the result through a web UI and REST API.

How it works

  1. Sources define where data comes from and which fields make a record unique
  2. Rules extract information with regex (extract or replace mode) — e.g. pull the merchant out of a transaction description
  3. Mappings turn extracted values into clean output — "DISCOUNT DRUG MART 32"{"vendor": "Discount Drug Mart", "category": "Healthcare"}
  4. Records are then queryable, pivotable, and exportable

Each record keeps three layers: data (raw import), transformed (rule and mapping output), and overrides (manual edits). Reads merge them in that order, so re-running the rules never clobbers something you typed by hand.

Stack

PostgreSQL with JSONB storage, a Node.js/Express API, and a React SPA served from public/. HTTP Basic auth, configured in .env.

Getting started

Requires PostgreSQL 12+, Node.js 18+, and Python 3.

npm install
python3 manage.py     # interactive setup: .env, database, schema, functions, UI, service

The UI is then at http://localhost:3020 and the API at http://localhost:3020/api (port set by API_PORT in .env).

For a walkthrough that creates a source, adds rules and mappings, and imports the sample CSV in examples/, see docs/getting-started.md.

Documentation

docs/getting-started.md Tutorial — build a working pipeline from scratch with curl
docs/spec.md Full reference — architecture, schema, data flow, API, manage.py
docs/ui.md Frontend: React + Vite build, key packages
docs/perspective.md Pivot table: pinned versions and API reference

Project structure

dataflow/
├── manage.py           # interactive setup / deploy / uninstall
├── database/           # schema.sql + one .sql file per API route
├── api/                # Express server, routes, auth middleware
├── ui/                 # React source (built to public/)
├── public/             # built UI, served as static files
├── docs/
└── examples/           # sample CSV for the tutorial

Both the API routes and the SQL are organized one file per resource, so api/routes/rules.js and database/rules.sql are the two halves of the same feature.

database/*.sql is the source of truth for every database function — never edit one directly in the database, or the next redeploy will silently revert it.

License

MIT