Commit Graph

12 Commits

Author SHA1 Message Date
e958bb5c6b Stop pretending Vercel is an option
Some checks failed
CI / Lint, Typen, Tests, Build (push) Failing after 5m51s
CI / Integrationstests (echtes Postgres) (push) Failing after 5m15s
It was never used. The repository lives on a self-hosted Gitea, which
Vercel's git integration cannot connect to at all — so the documented
route amounted to "mirror to GitHub first", and nobody did.

vercel.json is gone, and with it the branch in next.config.ts that
switched off `output: "standalone"` when the VERCEL variable was
present. That branch was the only functional trace; everything else was
documentation and comments describing a second deployment path that did
not exist.

DEPLOYMENT.md loses its "two supported ways" framing and the whole
Vercel section — about fifty lines. Several statements next to it were
stale for a different reason and are corrected in the same pass: the
outbound-firewall table still listed Supabase's pooler (the database is
a container now, nothing leaves the server), the prerequisites still
demanded an existing Supabase project, and the .env table still asked
for a pooler connection string instead of the two new passwords.

The nightly job is described as what it is — a container in
docker-compose.yml — rather than as a replacement for Vercel Cron.

Migrations keep their references: two comments from July mention Vercel
Cron, and they describe what was true when they were written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 11:02:40 +02:00
77d9a95f7f Run the database in a container of our own
Some checks failed
CI / Lint, Typen, Tests, Build (push) Failing after 5m25s
CI / Integrationstests (echtes Postgres) (push) Failing after 5m9s
Supabase was only ever the host: the application has talked to PostgreSQL
directly through pg/Kysely for a while. So the move is mostly about
supplying what the platform used to supply.

Proved before building anything. All 65 migrations replay onto an empty
database, and the result matches production exactly — 183 columns, 25
policies, 68 indexes, 84 constraints, identical sets, no diff. The only
function missing from the rebuild turned out to matter, see below.

What the platform supplied, deploy/db-init now does:

  - alpenwerk_app, explicitly NOBYPASSRLS. The whole access model is 21
    RLS policies; a role that bypasses them would leave everything
    working while showing too much, and nobody would notice.
  - pgcrypto and pg_trgm. uuid-ossp was available on Supabase but is
    used nowhere — no column default, no function calls uuid_generate_*.
  - anon, authenticated and service_role as NOLOGIN placeholders. No
    policy names them; they only carry grants the platform handed out,
    and a data dump referencing them would fail to restore without them.
  - A stub `auth` schema. The end state needs none of it — checked: no
    foreign key, no policy, no column default refers to it. The June
    2026 migrations do, and rewriting those would be falsifying history;
    they describe what was true then.

The gap the comparison found: rls_auto_enable() and the ensure_rls event
trigger existed only in the running database, created by hand, in no
migration. That is the net which forces RLS on every newly created
table — the reason a forgotten policy yields an empty table instead of
an open one. A rebuild from migrations would silently not have had it:
everything works, and the next new table is unprotected. Now a migration
(20260819100000), verified by creating a table on the rebuild and
confirming RLS came on by itself.

Data moves separately, via scripts/umzug-von-supabase.sh: schema from
the migrations, then pg_dump --data-only --disable-triggers for the rows.
Without --disable-triggers every foreign key trips over load order. RLS
does not interfere — none of the 19 tables uses FORCE ROW LEVEL
SECURITY, so the owner writes through. The dump is deliberately left on
disk afterwards.

psql and node come from two `tools`-profile services rather than being
installed on the host, so the server needs nothing but Docker. The db
service publishes no port at all — reachable only inside the compose
network.

SUPABASE_DB_URL is renamed MIGRATE_DATABASE_URL, since after this it
describes something else entirely; the old name still works so existing
.env files keep running. Both were exercised, as was the error when
neither is set.

The deploy workflow is set to manual-only. Its preconditions were never
met — no secrets, and whether the job container can reach the host's
Docker daemon is untested — and failing on every push teaches people to
ignore red runs. It also needs updating for the new database service
before it could work at all.

Not verified: none of this has run in an actual container. There is no
Docker daemon on this machine. What is verified is the part that
decides whether it can work — the schema, on a real empty database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 16:34:08 +02:00
d574d3c9d6 Aim the deploy at the runner that exists
Some checks failed
CI / Lint, Typen, Tests, Build (push) Failing after 5m28s
CI / Integrationstests (echtes Postgres) (push) Failing after 5m11s
Deploy / Migrationen und Container (push) Failing after 4s
The workflow asked for a `self-hosted` label. No runner on the instance
offers one, so the run would have sat in "Waiting" forever — no error, no
message, nothing to notice. The existing global runner elycon-runner-01
offers `docker` and `ubuntu-latest`, so both workflows now ask for
`ubuntu-latest`, the same label ci.yml already used.

That correction exposed a second thing the first version glossed over.
act_runner starts a container per job; mounting the Docker socket into
the *runner* does not put it in the *job*. Whether this job can reach the
host's daemon depends on the runner's config.yaml, which is not visible
from here — and the runner is global, so changing it affects every
repository on the instance, not just this one.

Rather than guess, the workflow now measures it in its first step and
fails with the fix if it cannot: which config lines to add for the
socket, or that SSH is the other way. Without that, the run would have
died three steps later on a message nobody could act on.

Both branches of the check were exercised: docker absent prints the
first message, docker present with no reachable daemon the second.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 16:15:49 +02:00
fc0989debb Deploy from a push, and start keeping track of migrations
Some checks failed
CI / Lint, Typen, Tests, Build (push) Failing after 5m52s
CI / Integrationstests (echtes Postgres) (push) Failing after 5m13s
Deploy / Migrationen und Container (push) Has been cancelled
A push to master now builds and restarts the application on the server:
a Gitea Actions workflow on a self-hosted runner writes .env from the
repository secrets, applies pending migrations, rebuilds the compose
stack against the host's Docker daemon, and waits for the container's
healthcheck before calling the run green. Without that last step a
deploy counts as successful the moment the container *starts*, even if
the app inside it dies immediately.

Switching migrations on automatically turned up something that had to be
fixed first: supabase_migrations.schema_migrations did not exist at all.
Every one of the 65 migrations was unrecorded, because they have been
applied by hand all along. An automatic `db push` would therefore have
replayed all 65 against the live database — initial_schema and the OM
cutover included. The database was checked against a spread of
migrations first (it is at head), then baselined: all 65 recorded as
applied without executing them.

The runner is scripts/migrate.mjs rather than the Supabase CLI. It needs
only `pg`, which the project already ships, instead of downloading a CLI
whose version drifts independently of this repository; and it does one
thing — the missing files, in order, each in its own transaction — where
`db push` also diffs schemas and may do more than that. Bookkeeping goes
in the same table in the same shape the CLI uses, so `supabase db push`
from a workstation still works and still skips what already ran.

The workflow lives in .github/workflows, not .gitea/. Gitea reads
.gitea/workflows and falls back to .github/workflows only when the
former is absent — creating .gitea/ would have silently switched off
ci.yml, with the run simply never appearing.

Verified: both workflow files parse; the secret check names what is
missing and refuses; values starting with "-" or containing "=" survive
being written to .env; and the runner was exercised against the real
database with a throwaway migration — applied once, skipped on a second
run, and on a deliberate syntax error rolled back whole, recording
nothing. Both probes were removed; the count is back to 65.

Not verified: nothing has run on an actual Gitea runner — none is
registered yet. DEPLOYMENT.md §5 covers registering one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 16:00:45 +02:00
d8a1fdf43b Reinstate the Vercel build settings
Reverts 61ccce5, which reverted ecbda3f. The decision came back to Vercel,
so the two platform accommodations return: output: "standalone" is
conditional on VERCEL again, and /api/import goes back to 60 seconds, the
free tier's ceiling.

The Docker path is unaffected and stays documented — including the internal
network notes and deploy/Caddyfile written in between, which remain correct
for anyone taking that road. DEPLOYMENT.md conflicted at the top and now
carries both introductions instead of one replacing the other.

Verified with VERCEL=1: builds clean and emits no standalone directory.

Stated once and recorded here rather than repeated: Vercel's Hobby plan
excludes commercial use, and this is a company's HR system. Defensible while
the database holds nothing but the 852 invented people from the seed;
Pro at $20/month is the licensed path once real personnel data is in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 08:32:47 +02:00
780f8fe2b7 Write down what an internal deployment actually needs
The target is a VM inside the company network, reachable only from there.
Two consequences decide whether this works at all, and both are easy to
discover too late — after the firewall rules are already written.

The server needs outbound access even though nothing comes in. Auth.js
exchanges the authorisation code for a token server-side and fetches the
issuer's configuration, so login.microsoftonline.com must be reachable from
the VM; the database likewise. That the person signs in through their own
browser is not enough, which is the assumption worth naming before someone
builds a closed network around it.

HTTPS is not optional either: Entra accepts http only for localhost. The
practical route without public reachability is a public DNS name pointing at
a private address and a certificate obtained through the DNS challenge —
allowed, common, and it yields a normally trusted certificate while the
server stays unreachable from outside. deploy/Caddyfile does that, and the
alternative (self-signed, trusted on every workstation) is written down with
its cost.

docker-compose now publishes port 3000 on 127.0.0.1 only. It was on every
interface, so the same service also stood there unencrypted, and one gap in
the firewall was enough. The proxy is the only way in.

AUTH_URL is documented for the same reason a comment sits in the Caddyfile:
behind a proxy the container does not see the name the browser used, and the
callback would point somewhere nobody can reach.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 08:26:14 +02:00
56662c0775 Describe what is actually deployed, now that the target is a Linux server
The revert restored three statements that stopped being true earlier today.

"Nicht containerisiert: Supabase (Datenbank + Auth)" — authentication is no
longer Supabase, it is Entra ID with an Auth.js session cookie, and the
database is any PostgreSQL 15 or later reached through DATABASE_URL. Supabase
is one option among several now, not the architecture.

The CI/CD note told the reader to pass --build-arg values for NEXT_PUBLIC_*.
Those variables no longer exist and the Dockerfile stopped taking build
arguments today. Following it would produce a puzzling failure; the point
now is the opposite one, that no build arguments are needed at all and the
same image runs everywhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 08:23:30 +02:00
61ccce5456 Revert "Make the build fit Vercel without breaking the container"
This reverts commit ecbda3f. The deployment goes to a Linux server instead,
so the two accommodations no longer earn their place: output: "standalone"
returns to unconditional, which is what the Dockerfile wants, and
/api/import goes back to 120 seconds — the free-tier ceiling that forced 60
does not apply outside a serverless platform, and a large import benefits
from the headroom.

The Vercel section in DEPLOYMENT.md goes with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 08:21:58 +02:00
ecbda3f3a5 Make the build fit Vercel without breaking the container
Two settings were wrong for a platform build.

output: "standalone" tells Next.js to emit a self-contained server, which is
what the Dockerfile copies in — and what Vercel neither needs nor expects,
since it builds and packages the app itself. It is now conditional on the
VERCEL variable, which every build there sets, so each path gets what it
wants. Verified both ways: with VERCEL=1 no standalone directory appears,
without it one does.

/api/import declared maxDuration = 120. The free tier caps at 60 and refuses
anything higher, so the deployment would have failed on a value chosen for a
self-hosted server. Lowered, with the reason and the Pro ceiling written
next to it.

DEPLOYMENT.md now covers both paths, and says plainly that the repository
cannot be connected: git.elycon.solutions is self-hosted, and Vercel's git
integration only speaks GitHub, GitLab and Bitbucket. Deploying from the
workstation with the CLI works with any repository and is the shorter road;
mirroring to GitHub is written down as the alternative, with its cost — two
remotes to keep in step.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 15:17:24 +02:00
99e50fbbf9 Stop hoarding connections the provider will not give twice
The app died with "max clients reached in session mode - pool_size: 15".
Two causes, both real, neither visible without a live database.

The connection string pointed at the pooler's session mode, which pins one
backend per client and caps at 15 on Supabase. Every query here already runs
inside a transaction and the session context is set transaction-locally, so
transaction mode is not a workaround but the mode this design was written
for. Verified: 20 concurrent transactions, all 852 rows, 0.4s — and still
nothing without a session context.

The second cause was the dev server. Next.js re-evaluates changed modules,
so a module-local `let` was empty afterwards while the previous pool stayed
alive holding its connections. An afternoon of editing exhausted the quota.
The pool now hangs off globalThis, which is inert in production where
nothing reloads.

Documented in .env.example and DEPLOYMENT.md, because a deployment that
picks port 5432 fails this way under load and not before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 10:28:36 +02:00
2ba9b37aa7 Hand the front door to Entra, and keep the keys out of the build
Auth.js replaces GoTrue. The sign-in still goes to the same Entra tenant,
but nothing sits between the app and the identity provider any more — the
code exchange, state, nonce and the session cookie are ours.

lib/auth/session.ts stays the only place that knows where a user id comes
from, which is why this was one file and not fifty. What it returns is now
app_users.id. app_upsert_user() maps the Entra `oid` onto it, and for an
address that already has a profiles row it adopts that id instead of
minting a new one — otherwise everyone would have been signed in and cut
off from their own notes, drafts and audit trail at the same time.

That upsert is the one write that cannot have a session context yet: the
id is what it produces. It runs as a SECURITY DEFINER function that may
touch app_users and nothing else, which is a far smaller lever than the
service key that used to answer this class of problem.

The proxy no longer checks HR rights. It has no database connection, and
putting role/is_active in the token would have frozen the claim until the
next sign-in. The check moved to where it can read the current truth: the
app layout on every render, requireHrUser() for the export routes, and
underneath both, RLS.

Two things only came out by running it:

  - `export const proxy = auth(…)` is not a function declaration, so
    Next.js never found it and every request 404'd. `next build` reported
    success and listed the proxy. In the function config form auth() also
    returns the handler as a promise, so it needs an await. The proxy test
    now mocks it as a promise for that reason — a friendlier mock would
    let the same bug back in.

  - A missing AUTH_MICROSOFT_ENTRA_ID_ISSUER silently falls back to
    /common/, and the redirect really did go there. That would let any
    Microsoft account sign in, including a private one, and it would never
    look broken. It now refuses to start in production.

Neither build nor image needs credentials any more: the pool is created on
first use, the auth config is evaluated per request, and there are no
NEXT_PUBLIC_* values left to bake in. One image now runs in every
environment.

Verified: typecheck, lint, 187 tests, build, and by hand in the browser —
/employees redirects to /login, and the sign-in button reaches the Entra
page with PKCE and the callback URL that goes into the app registration.
Not verified against a real database; there is still no DATABASE_URL.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 14:57:32 +02:00
79f0e19bf8 Org assignment history, mobile support, and a correctness pass
Data model
- employee_assignments records org placement over time (valid_from/valid_to),
  written by a trigger on `employees` rather than inside each RPC: ~70
  `update employees` statements spread over fifteen migrations mean per-call
  bookkeeping would miss paths today and again with every future RPC. A
  partial unique index enforces the one-open-interval invariant the trigger
  relies on when closing the current row.
- The Organigramm gains a Stichtag (default today). Membership comes from
  entry/exit/karenz, past placement from the new history, future placement
  projected from pending_org_changes. Placements predating the migration are
  backfilled with today's values and flagged as such in the UI, since
  employee_history only ever stored free text and cannot be reconstructed.

Correctness
- Reports and exports silently truncated at PostgREST's 1000-row cap
  (db.max_rows); employee_history is already past it at ~800 staff. Every
  whole-table read now pages explicitly.
- XLSX date cells were a day early: ExcelJS converts a Date to an Excel
  serial straight off getTime(), so a Date built at local midnight lands on
  the previous day's serial in any positive-offset zone.
- Date handling is pinned to Europe/Vienna throughout, and date-only strings
  are formatted without a Date round-trip. The dashboard's YTD window was
  built by round-tripping a local Date through toISOString(), which shifted
  it a day early and dropped 31 December entirely.
- Export routes parsed measure/group/split/eventType with unchecked `as`
  casts, so an unknown value reached column headers as `undefined` and the
  Content-Disposition filename. Parsed against the label maps now, with the
  filename slugged as a backstop.
- toXlsx keyed columns by header text, silently dropping the second of any
  two columns sharing a name — split columns take their header from data.
- The org chart tree walks had no cycle guard; nothing in the schema forbids
  a manager_id cycle, and one would hang the tab rather than misreport.
- The login page reflected ?error= verbatim, letting anyone put arbitrary
  text on the real sign-in screen; messages are looked up by code now.
- React Flow needs elementsSelectable on, or it sets pointer-events:none on
  the whole node and the expand control stops responding.

UI
- Mobile: the shell was unusable below lg — a fixed 236px margin pushed
  content off-screen with no mobile navigation at all. The sidebar is now a
  drawer, dvh replaces vh, safe-area insets are honoured, inputs are 16px so
  iOS stops zooming on focus, and form grids stack.
- Org chart nodes redesigned: per-kind accent stripes and icons, vacant
  roles called out, expand control moved to the bottom edge carrying the
  child count.
- Pagination is windowed; it previously rendered one link per page (54 for
  the employee list, unbounded for the audit log).
- Positions page reduced to open positions with a single "Besetzen" action.
- The employee Organisation tab links into the org chart focused on that
  person, reusing the chart's existing search-match highlighting.

Also included, uncommitted until now
- Dependants, HR notes, academic titles, split address fields, position
  validity and role/employment fields, with their migrations and UI.
- Docker/compose deployment setup, data-model and security-review docs.
2026-07-24 23:38:10 +02:00