The app died with "max clients reached in session mode - pool_size: 15".
Two causes, both real, neither visible without a live database.
The connection string pointed at the pooler's session mode, which pins one
backend per client and caps at 15 on Supabase. Every query here already runs
inside a transaction and the session context is set transaction-locally, so
transaction mode is not a workaround but the mode this design was written
for. Verified: 20 concurrent transactions, all 852 rows, 0.4s — and still
nothing without a session context.
The second cause was the dev server. Next.js re-evaluates changed modules,
so a module-local `let` was empty afterwards while the previous pool stayed
alive holding its connections. An afternoon of editing exhausted the quota.
The pool now hangs off globalThis, which is inert in production where
nothing reloads.
Documented in .env.example and DEPLOYMENT.md, because a deployment that
picks port 5432 fails this way under load and not before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Auth.js replaces GoTrue. The sign-in still goes to the same Entra tenant,
but nothing sits between the app and the identity provider any more — the
code exchange, state, nonce and the session cookie are ours.
lib/auth/session.ts stays the only place that knows where a user id comes
from, which is why this was one file and not fifty. What it returns is now
app_users.id. app_upsert_user() maps the Entra `oid` onto it, and for an
address that already has a profiles row it adopts that id instead of
minting a new one — otherwise everyone would have been signed in and cut
off from their own notes, drafts and audit trail at the same time.
That upsert is the one write that cannot have a session context yet: the
id is what it produces. It runs as a SECURITY DEFINER function that may
touch app_users and nothing else, which is a far smaller lever than the
service key that used to answer this class of problem.
The proxy no longer checks HR rights. It has no database connection, and
putting role/is_active in the token would have frozen the claim until the
next sign-in. The check moved to where it can read the current truth: the
app layout on every render, requireHrUser() for the export routes, and
underneath both, RLS.
Two things only came out by running it:
- `export const proxy = auth(…)` is not a function declaration, so
Next.js never found it and every request 404'd. `next build` reported
success and listed the proxy. In the function config form auth() also
returns the handler as a promise, so it needs an await. The proxy test
now mocks it as a promise for that reason — a friendlier mock would
let the same bug back in.
- A missing AUTH_MICROSOFT_ENTRA_ID_ISSUER silently falls back to
/common/, and the redirect really did go there. That would let any
Microsoft account sign in, including a private one, and it would never
look broken. It now refuses to start in production.
Neither build nor image needs credentials any more: the pool is created on
first use, the auth config is evaluated per request, and there are no
NEXT_PUBLIC_* values left to bake in. One image now runs in every
environment.
Verified: typecheck, lint, 187 tests, build, and by hand in the browser —
/employees redirects to /login, and the sign-in button reaches the Entra
page with PKCE and the callback URL that goes into the app registration.
Not verified against a real database; there is still no DATABASE_URL.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Erster Schritt weg von Supabase hin zu "läuft auf jedem PostgreSQL".
Gemessen sitzt die Kopplung nicht dort, wo der Begriff "Supabase-Projekt"
sie vermuten lässt: das Schema ist reines PostgreSQL, und von 58 RLS-Policies
rufen nur fünf auth.uid() direkt auf. Die übrigen 53 gehen über is_hr_user().
Diese eine Funktion ist die Brücke — wird sie umgelegt, folgt der Rest.
Die Migration legt sie um. app_current_user_id() liest jetzt zuerst
current_setting('app.user_id') und fällt nur ersatzweise auf auth.uid()
zurück. Deshalb plpgsql statt language sql: eine SQL-Funktion wird beim
Anlegen geparst, und auth.uid() gibt es auf einem gewöhnlichen PostgreSQL
nicht — die Migration liesse sich dort gar nicht erst anwenden. Der
Ausnahmeblock fängt das ab, und damit läuft dieselbe Migration auf beiden
Systemen. Der Rückfall verschwindet mit der Abschlussmigration.
Dazu app_users als Nachfolger von auth.users, external_id ist die oid des
Anbieters statt der E-Mail: eine Namensänderung darf kein zweites Konto
erzeugen.
Die neue Zugriffsschicht ist Kysely auf einem pg-Pool. Was daran zählt, ist
nicht der Query-Builder, sondern was er verhindert:
- Die Kysely-Instanz wird nicht exportiert. Wer abfragen will, geht durch
withUser() — und das öffnet immer eine Transaktion.
- set_config(..., true) ist transaktionslokal. Ohne das dritte Argument
bliebe die Kennung an der gepoolten Verbindung kleben und die nächste
Anfrage liefe im Namen der vorherigen Person. In einer Personaldatenbank.
- Eine ESLint-Regel verbietet den Import von pg und von lib/db/pool
ausserhalb von lib/db. Nachgewiesen: eine Testdatei mit beiden Importen
erzeugt zwei Fehler.
- Einen privilegierten Zugang gibt es nicht mehr. asSystem() benutzt
dieselbe Rolle ohne BYPASSRLS; was ohne angemeldete Person laufen darf,
muss als SECURITY-DEFINER-Funktion in der Datenbank stehen.
tests/integration/session-context.test.ts läuft gegen einen Pool mit genau
einer Verbindung — sonst träfe er die Lücke mal und mal nicht. Er prüft, dass
nach Commit *und* nach Rollback nichts an der Verbindung zurückbleibt, und
belegt in einer Gegenprobe, dass eine Einstellung ohne Transaktion tatsächlich
hängen bleibt. Ein Sicherheitstest, der sich mangels DATABASE_URL selbst
überspringt, wäre schlimmer als keiner: in der CI schlägt schon das Fehlen
des Verbindungsstrings fehl.
Beim Schreiben der Migration stellte sich heraus, dass die Policies
hire_drafts_owner und saved_reports_owner heissen, nicht _own. Mit dem
geratenen Namen hätte drop policy nichts getroffen und create policy wäre mit
"already exists" abgebrochen.
Typecheck, Lint und 182 Tests sind grün. Die Anwendung läuft unverändert
weiter — sie benutzt die neue Schicht noch nicht.
Data model
- employee_assignments records org placement over time (valid_from/valid_to),
written by a trigger on `employees` rather than inside each RPC: ~70
`update employees` statements spread over fifteen migrations mean per-call
bookkeeping would miss paths today and again with every future RPC. A
partial unique index enforces the one-open-interval invariant the trigger
relies on when closing the current row.
- The Organigramm gains a Stichtag (default today). Membership comes from
entry/exit/karenz, past placement from the new history, future placement
projected from pending_org_changes. Placements predating the migration are
backfilled with today's values and flagged as such in the UI, since
employee_history only ever stored free text and cannot be reconstructed.
Correctness
- Reports and exports silently truncated at PostgREST's 1000-row cap
(db.max_rows); employee_history is already past it at ~800 staff. Every
whole-table read now pages explicitly.
- XLSX date cells were a day early: ExcelJS converts a Date to an Excel
serial straight off getTime(), so a Date built at local midnight lands on
the previous day's serial in any positive-offset zone.
- Date handling is pinned to Europe/Vienna throughout, and date-only strings
are formatted without a Date round-trip. The dashboard's YTD window was
built by round-tripping a local Date through toISOString(), which shifted
it a day early and dropped 31 December entirely.
- Export routes parsed measure/group/split/eventType with unchecked `as`
casts, so an unknown value reached column headers as `undefined` and the
Content-Disposition filename. Parsed against the label maps now, with the
filename slugged as a backstop.
- toXlsx keyed columns by header text, silently dropping the second of any
two columns sharing a name — split columns take their header from data.
- The org chart tree walks had no cycle guard; nothing in the schema forbids
a manager_id cycle, and one would hang the tab rather than misreport.
- The login page reflected ?error= verbatim, letting anyone put arbitrary
text on the real sign-in screen; messages are looked up by code now.
- React Flow needs elementsSelectable on, or it sets pointer-events:none on
the whole node and the expand control stops responding.
UI
- Mobile: the shell was unusable below lg — a fixed 236px margin pushed
content off-screen with no mobile navigation at all. The sidebar is now a
drawer, dvh replaces vh, safe-area insets are honoured, inputs are 16px so
iOS stops zooming on focus, and form grids stack.
- Org chart nodes redesigned: per-kind accent stripes and icons, vacant
roles called out, expand control moved to the bottom edge carrying the
child count.
- Pagination is windowed; it previously rendered one link per page (54 for
the employee list, unbounded for the audit log).
- Positions page reduced to open positions with a single "Besetzen" action.
- The employee Organisation tab links into the org chart focused on that
person, reusing the chart's existing search-match highlighting.
Also included, uncommitted until now
- Dependants, HR notes, academic titles, split address fields, position
validity and role/employment fields, with their migrations and UI.
- Docker/compose deployment setup, data-model and security-review docs.
Reworks the app from a two-role (hr_admin/manager) model to a single
HR-only role gated by profiles.is_active, fixes transfer/promote/karenz/
reorg RPCs to actually defer future-dated changes via a new
pending_org_changes table instead of writing them immediately (applied
by a daily Vercel Cron route), makes reorg undo append-only instead of
deleting history, adds Karenz-return and history-date integrity guards,
deprecates the salary column, and adds explicit schema grants + perf
indexes needed to run against a fresh (non-hosted) Postgres instance.
Adds vitest unit + integration test suites (the latter against a real
local Supabase instance) covering all of the above, plus lint/typecheck/
build wiring (`npm run check`).