pf_app/CLAUDE.md
Paul Trowbridge 4d2af589b9 Document both axes collapsing through ViewConfig
§Row depth and the observer shim and §Column hierarchy described two
different mechanisms and an asymmetry between the axes that no longer
exists: rows were imperative view state chased by an observer shim, and
columns collapsed by truncating split_by. Both are a depth in ViewConfig
now, so the two sections are one.

Records what matters for reading the code: that the depths are 1-based
against server.cpp's `set_depth(depth - 1)`, that this needed the engine
patch in ui/vendor because apply_update dropped the fields, and that
Session::update_view_config returning early on "no change" is why nothing
happened at all rather than happening wrongly.

Keeps a short account of what was deleted -- the shim, the focus listeners,
the retry loop, splitFull, collapsingRef, the prefix test -- because someone
will otherwise wonder why the surrounding code looks so plain, and because
split_full is still read for layouts saved under the old scheme.

Per-node expansion moves to its own subsection with the reason it is hard
(set-only API, no getter) and what fixing it would take. The stale claim that
the selection is cleared on every column change is gone: with a depth the
dimensions a slice was cut from are all still present.

Also: observerShim.js dropped from the project layout, and the tech stack
now says 5.4.0 from the patched vendored build rather than 5.2.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 12:38:18 -04:00

21 KiB
Raw Blame History

Pivot Forecast — CLAUDE.md

What this app is

A web app for building named forecast scenarios against any PostgreSQL table. The workflow: load historical actuals as a baseline (optionally date-shifted into the forecast period), then apply incremental adjustments (scale, recode, clone) to build a plan. All changes are append-only, fully audited, and reversible by log entry.

Full spec: pf_spec.md Data transport architecture options: pf_perspective_options.md


Tech stack

  • Backend: Node.js / Express (server.js)
  • Database: PostgreSQL — isolated pf schema
  • Frontend: React + Vite + Tailwind CSS in ui/; built output lands in public/app/
  • Pivot: Perspective (@perspective-dev/* distribution, not FINOS @finos/perspective) 5.4.0 from a patched build vendored in ui/vendor — see its README for what is patched and which host-tool versions the rebuild needs; bundled inline via the /inline entrypoints — never from a CDN (the 4.x CDN bundle resolves its server WASM to an unversioned path and silently pulls whatever is newest). See PERSPECTIVE.md.
  • Dev: npm run dev (nodemon) in root; npm run build in ui/

Project layout

server.js               Express entry point; pg pool; session; type parsers for bigint/numeric
routes/
  auth.js               POST /api/login, /api/logout, GET /api/me; login throttle
  tables.js             GET /api/tables, /api/tables/:schema/:tname/preview
  sources.js            Source registration, col_meta, SQL generation
  versions.js           Version CRUD, baseline/reference load, data stream
  operations.js         scale, recode, clone, undo — the core forecast ops
  log.js                GET /api/versions/:id/log, DELETE /api/log/:logid
lib/
  sql_generator.js      buildFilterClause, token substitution helpers
  auth.js               scrypt hash/verify, requireAuth, sessionUser; `node lib/auth.js hash` CLI
  utils.js
setup_sql/
  01_schema.sql         pf schema DDL — run once to install
  02_auth.sql           pf.app_user + pf.session
ui/src/
  auth.jsx              AuthProvider/useAuth; wraps fetch so any 401 returns to login
  views/
    Login.jsx           Sign-in form
    Setup.jsx           DB browser, source registration, col_meta editor
    Baseline.jsx        Version management, baseline workbench, reference load
    Forecast.jsx        Perspective pivot, selection handling, operation dispatch
  components/
    OperationPanel.jsx  The adjustment workbench — ledger + scale/recode/clone forms
    BridgeView.jsx      Baseline → current waterfall by tag (exports buildSteps/layoutSteps)
    Sidebar.jsx         3-step collapsible nav
    StatusBar.jsx       Source · version · write target · row counts · theme
    Timeline.jsx        Date-range preview bar for baseline segments

Database schema (pf)

  • pf.source — registered source tables
  • pf.col_meta — column roles: dimension | value | units | date | filter | ignore; dim_group groups functionally dependent columns (e.g. a part and its attributes, or a date and its derived year/month dimensions); dim_period_col maps a dimension to a pf.dim_period column so date-adjacent values are derived at load time rather than copied raw; in_grain flags dimension/date columns that define the display grain (see below); is_key is described under §is_key below
  • pf.version — named forecast scenarios; exclude_iters (default ["reference"]) blocks those iter values from all operations
  • pf.fc_{tname}_{version_id} — one forecast table per version; contains both operational rows (pf_iter = baseline|scale|recode|clone) and reference rows (pf_iter = reference). Indexed on pf_logid, which is how undo, the change-log aggregate and the grain key all find their rows — without it each is a sequential scan of the whole table. Tables created before that index was added do not have it.
  • pf.log — audit log; every write gets one entry; slice + params stored as jsonb
  • pf.sql — generated SQL templates per source/operation; tokens substituted at request time
  • pf.app_user — login accounts; scrypt pass_hash, is_active, last_login_at
  • pf.session — express-session store (connect-pg-simple layout)
  • pf.dim_period — calendar lookup table (20182035); one row per month keyed on sdat (month start date); provides cal/fiscal year, quarter, and month columns; populated by setup_sql/gen_dim_period.sql with a configurable fiscal year start month

is_key

Read in four places, and not the one it sounds like — slices are validated against filterCols, which is every dimension plus every date column regardless of is_key.

  1. The key of a dim_groupresolveGroup() in routes/sources.js takes members.find(c => c.is_key), the column every other member is keyed on for pf.dim_member
  2. Value completion — only is_key columns get a dropdown, and /sources/:id/values/:col refuses anything else
  3. Sibling autofill — fires on blur only when is_key && dim_group
  4. The dim_period anchorrole === 'date' && is_key && dim_group picks the date whose siblings are derived from the calendar

Uses 2 and 3 want several columns flagged; use 1 needs exactly one per group. When a group has more than one, .find() silently takes the lowest opos. That is not hypothetical: segment_new (opos 12) outranked part (opos 15) in the part group, so a refresh keyed on a column that is null throughout, matched nothing, and reported success with zero members. customer's group has four keys and picks the right one only by opos luck.

The two meanings want separating — a per-group key choice, or a rule that the group key is the column named by the group (which these groups nearly follow already, except sdatesdate_e). Until then, a refresh that finds two keys should refuse rather than guess.

Key token substitution tokens

{{fc_table}}, {{where_clause}}, {{exclude_clause}}, {{logid}}, {{pf_user}}, {{value_incr}}, {{units_incr}}, {{pct}}, {{set_clause}}, {{scale_factor}}, {{date_offset}}, {{filter_clause}}


Core data flow

Initial load (Forecast view)

Forecast.jsx fetches col_meta first, then picks the endpoint:

  • grain mode (any in_grain column) — GET /api/versions/:id/agg, rows pre-aggregated to the grain, table indexed on pf_gkey
  • raw mode (no grain) — GET /api/versions/:id/data, raw forecast rows, table indexed on pf_id

Either way: Arrow IPC binary stream → worker.table(buffer) in Perspective WASM. fetchArrow() handles both.

Why one batch (not streaming): pg returns bigint/numeric as strings by default — type parsers in server.js coerce them to numbers. Per-batch Arrow encoding creates independent dictionaries that cause Perspective WASM to crash on dictionary replacement messages. Server accumulates all rows, emits one record batch.

Display grain

Aggregating to the grain the pivot actually displays is the load-time fix — measured 534,902 → 6,154 rows on osm_stack. It keeps the native Perspective engine, so expand/collapse/depth/sort/filter all still work. Set the grain in Setup (in_grain per column); it is baked into pf.sql at Generate SQL time so load and operations agree. grainOf() in lib/sql_generator.js is the single definition of what the grain is — Setup.jsx and routes/log.js mirror it. Full design: pf_spec.md → §Display-grain pre-aggregation. Why not a DuckDB virtual server: pf_perspective_options.md → §Spike findings.

Segment and note columns

/data and /agg both LEFT JOIN pf.log and emit two columns the forecast table does not itself carry:

  • pf_segment — for a baseline or reference row, that load's label (tag, else note); '(adjustment)' for everything else
  • pf_note — the free text on a scale/recode/clone; null on loads

They are deliberately separate: commingling a segment name with an adjustment note makes neither pivotable. In grain mode pf_logid is part of the grain, so the join adds no rows. The operation routes stamp the same two fields onto the rows they push back incrementally, since those come from RETURNING * and would otherwise arrive without them.

Forecast operations

POST to /api/versions/:id/{scale|recode|clone} → SQL executed with RETURNING * → new rows returned as JSON → pspTable.update(rows) — no full reload. In grain mode the operation's final CTE aggregates its own new rows to grain first; since pf_logid is part of pf_gkey those keys are always new, so update() appends and the view re-sums.

Price or volume (plug)

A sales figure alone does not say which of price or volume moved, so scale takes plug'price' (default, volume holds) or 'volume' (price holds, units scale in proportion). Resolved in resolveIncrs() in routes/operations.js; the panel only offers it when the edit is dollars-only, because naming units or price has already answered it. The semantics come from the predecessor Excel model, /opt/forecast_api/VBA/fpvt.frmcalc_val / calc_price:

plug volume:  pchange = fVal/(pVal+bVal); fVol = (pVol+bVol)*pchange
plug price:   fVol = pVol + bVol

A target_price with a target_units alongside is that form's Edit Price mode, where both are inputs and dollars fall out. The ledger's Result line previews value, units and price together using the same rules, so what you see is what gets written.

Undo

DELETE /api/log/:logid → removes rows by logid → table.remove() of the affected index values (pf_gkeys in grain mode, pf_ids in raw mode); the view re-sums. No full reload.


Axis depth (collapse / expand)

Both axes collapse the same way: a depth in ViewConfig, set through restore().

  • Rowsgroup_by_depth, driven by the EXPAND 0 1 2 3 buttons via applyDepth()
  • Columnssplit_by_depth, driven by the COLUMNS buttons via applySplitDepth()

Both are 1-based: they count the levels to show, where the imperative view.set_depth() counts the boundary below them. server.cpp does ctx1->set_depth(row_pivot_depth - 1) and ctx2->set_depth(HEADER_COLUMN, column_pivot_depth - 1), so the toolbar sends d + 1 and subtracts one again when reading a layout back.

Because a depth lives in the config, it survives every view rebuild, rides into the persisted and named layouts through viewer.save(), and needs nothing re-applied afterwards.

This required patching the engine. ViewConfig::apply_update applied ten fields and neither depth, so a depth sent through restore() was accepted, deserialized and dropped — and since Session::update_view_config returns early when apply_update reports no change, no view was rebuilt at all. See ui/vendor/0001-apply-depth-fields-on-config-update.patch; the vendored build carries it.

What this replaced

Worth knowing, because a lot of machinery existed to work around it and is now gone:

  • Row depth used to be imperative — getView() then view.set_depth() — which put it on an object the viewer discards whenever it re-renders. Restoring it meant guessing when that had happened: an observerShim.js patching window.IntersectionObserver and window.ResizeObserver, focus and visibility listeners, a retry loop for getView() throwing No table set while getTable() resolved, and a flag tracking whether the viewer had "gone away". That guesswork produced three distinct visible faults — the tree fully expanding on refocus, snapping on any reflow, and snapping when Perspective's settings sidebar opened.
  • Column collapse used to restore a truncated split_by. The discarded levels therefore had to be remembered separately (splitFull, persisted as split_full), our own collapse had to be told apart from the user rearranging the pivot (collapsingRef plus a prefix test), and the selection was cleared on every collapse because the axis was changing shape.

split_full is still read on load, for layouts saved under the old scheme.

Still not solved: per-node expansion

Expanding one specific branch is view state with no config representation, and the API has no getter:

expand(row_index: number): Promise<number>
collapse(row_index: number): Promise<number>

It can be set but not read, so it cannot be captured and replayed — this is why it is lost on every rebuild, and why no client-side fix has worked. The honest route is a ViewConfig field carrying expanded row paths (indices shift as the tree opens), applied in server.cpp where the depths are. Bigger than the depth patch: it needs a way to enumerate expanded nodes in the C++ traversal, a proto field, and the apply step.

The fork's own header_click.ts / expand_column / collapse_column is the column-axis equivalent and has the same limitation.

Limitation that remains either way: depth is whole-axis. Excel can collapse 2025 while 2026 stays expanded; a depth collapses every group at that level together. columns selects which measures appear, not individual split combinations.

Slice mechanics

When the user clicks a pivot cell, perspective-click fires. The handler in Forecast.jsx extracts [col, '==', value] filters from detail.config.filter — only role = dimension and role = date columns are kept as the slice. A plain click replaces the selection; ctrl/⌘/shift-click toggles a slice in or out of it, so the panel holds a list of slices sent as slices in operation POST bodies (the single slice object is still accepted server-side).

Dragging across a block of cells selects a region. The datagrid runs in edit_mode: SELECT_REGION (forced on restore, so a saved layout can't switch it off) and reports the region as a perspective-select event carrying a Perspective ViewWindow{ start_row, end_row, start_col, end_col }, not the per-row insertConfigs payload an older API used. It fires on every mouseover as the region grows, so the handler only records the latest window and a window-level mouseup commits it. A single-cell region is ignored there: perspective-click already owns plain and modifier clicks, and handling it in both places would undo a ctrl-click toggle.

Turning a region back into slices re-derives, per cell, the same filters Perspective attaches to a click — row dimensions from the view's __ROW_PATH__ (raw values, so dates stay epoch millis rather than whatever the grid formatted them as), column dimensions from the split_by segments of the column name. The grand-total row resolves to no dimension at all and is skipped; that would mean "the whole version".

Selection highlight. The datagrid highlights whatever sits in its own model._selection_state.selected_areas, and wipes that list on every mousedown — so a multi-slice selection built up over several ctrl-clicks would only ever show the last cell. Forecast.jsx keeps areasRef, a sliceKey -> rectangles map parallel to slices, and an effect pushes the full set back and redraws after every change. Deselecting anywhere (ctrl-click, the panel's ×, Clear selection) prunes the map by live slice key, so the grid and the panel can't disagree.

pf_iter is not a col_meta column, so it is stripped when a slice is built: two cells differing only by iter band produce the same effective slice. Duplicates are collapsed before the request — without that, apply_mode: each would apply the same change twice.

Limitation: computed columns created by Perspective's split_by (e.g. Month, YearDate) don't map back to raw rows — only native dimension columns work for slice extraction.


Operation SQL patterns

All three operations follow the same structure: insert a pf.log row in a CTE, then insert forecast rows referencing its id. {{where_clause}} is built from the slice; {{exclude_clause}} blocks exclude_iters rows.

  • Scale — distributes value_incr/units_incr proportionally across rows in the slice using window functions
  • Recode — inserts negative rows (zero out original) + positive rows with {{set_clause}} dimension overrides; both share the same logid
  • Clone — copies the slice with {{set_clause}} overrides and {{scale_factor}} multiplier; original untouched

build_where() validates every slice key against col_meta (only role = dimension allowed). Values are escaped but not parameterized — consistent with existing patterns, debuggable in pg logs.


Authentication

Everything under /api except the auth routes sits behind a session; the React app is only mounted once there is one (Gate in main.jsx), because its load effects call the API immediately.

  • Accounts: pf.app_user — scrypt hashes from lib/auth.js, never plaintext. Managed with ./pf.sh add-user | passwd | list-users | disable-user | enable-user; the password is read on stdin and hashed before it reaches psql.
  • Sessions: express-session + connect-pg-simple in pf.session, so a restart doesn't sign everyone out and a session can be revoked by deleting its row (disable-user does exactly that). Cookie pf.sid: httpOnly, SameSite=Lax, Secure unless COOKIE_SECURE=false, 12h rolling.
  • Config: SESSION_SECRET is required — the server exits at boot without one. TRUST_PROXY (default 1) makes req.ip and secure-cookie detection correct behind the TLS proxy. CORS_ORIGIN is the only way CORS is enabled at all; a wildcard origin plus a session cookie would be cross-site request forgery by construction.
  • Login hardening: routes/auth.js throttles to 10 failures per IP per 15 minutes (in-memory), returns one message for unknown/wrong/disabled alike, and regenerates the session id on success.

Identity is server-side. pf_user, created_by and closed_by come from sessionUser(req), never from the request body — the UI used to send a hardcoded pf_user: 'admin', which any client could have set to anything. The audit log now names the account that made the change.

Light / dark mode

Theme state lives in ui/src/theme.jsx — a React context (ThemeContext) with a ThemeProvider that wraps the app in main.jsx.

  • Storage key: pf_dark in localStorage; falls back to window.matchMedia('(prefers-color-scheme: dark)') on first visit
  • Toggle: setDark(d => !d) in StatusBar.jsx; effect writes localStorage and toggles the .dark class on <html>
  • CSS: ui/src/index.css defines CSS custom properties under :root (light) and .dark. All Tailwind color overrides are written as .dark .bg-white { ... } etc. — no Tailwind dark-mode config needed
  • Palette: dark mode uses Perspective's "Pro Dark" colours (--bg-primary: #242526, panels #2a2c2f, gridlines #3b3f46, text #c5c9d0)
  • Perspective viewer: Forecast.jsx calls viewer.setAttribute('theme', dark ? 'Pro Dark' : 'Pro Light') both on initial load and in a useEffect([dark, versionId]) so the viewer stays in sync when the toggle fires
  • Consuming the theme: import useTheme from '../theme.jsx' then const { dark, setDark } = useTheme()

Known issues / active work

  • Load time is dominated by row count, not payload size. On fc_osm_skinny_29 (2.56M raw rows) a 24-column grain still yields 285,685 rows: ~2s to aggregate in pg, but ~15s to serialise those rows out of Postgres and parse them into JS, then ~3s to build Arrow. Halving the payload (the pf_gkey md5) barely moved it. The remaining lever is dynamic grain — group by the fields the current pivot actually uses rather than every in_grain column; see pf_spec.md → §Display-grain pre-aggregation, "the dynamic variant"
  • Per-node expand/collapse is lost whenever the view rebuilds — set-only API, no getter; see §Axis depth
  • Default pivot layout should be configurable per source (currently hardcodes first 2 dimensions)
  • Source/version selection persists in localStorage (pf_sourceId / pf_versionId, App.jsx). It is re-validated against the live list whenever that list changes, so a deregistered source or deleted version re-points at the first remaining one instead of leaving a dead id that 404s every call
  • Col_meta / version schema drift: if col_meta roles change after a version's forecast table is created, SQL and DDL go out of sync — workaround is to delete and recreate the version
  • Grain drift: changing in_grain after a load requires Generate SQL + a page reload, since the loaded table's index and columns are fixed at load time. routes/log.js derives the grain from live col_meta, so a grain changed mid-session yields pf_gkeys that don't match the loaded table and undo silently removes nothing
  • Grain is static per source — a dimension left unflagged cannot be pivoted on. Dynamic per-cut grain (intersect the viewer's field set with the eligible set) is the additive next step; see pf_spec.md → §Display-grain pre-aggregation

Deferred (not in v1)

Baseline replay (replay: true returns 501), approval workflow, territory filtering, export, version comparison, multi-DB connections. Live server-side aggregation (Path A / DuckDB virtual server) is parked on branch spike/duckdb-virtual-server.