Match search words at the start of a word, not anywhere inside one

Searching "Winkler H" returned all seven Winklers instead of the one
Hannah. Each word was matched as a substring, so "H" hit T-h-omas,
Kat-h-arina and CNC-Dre-h-er:in — every row. The shorter the input, the
more useless the result, and an initial is the shortest input anyone
would type.

A word now has to match at the start of a word: either the haystack
begins with it, or a space does. The haystack is first name, last name
and job title joined, with hyphens, slashes, colons and dots flattened
to spaces, so "dreher" still finds CNC-Dreher:in and "cnc" still finds
both the Dreher and the Fräser.

Checked against the live data before and after: "winkler h" now returns
Hannah Winkler alone, "h winkler" the same in either order, "winkler
kat" the two Katharinas, "dreher" the twelve CNC-Dreher.

The trigram index on the concatenated name no longer applies, which is
the price. At under nine hundred rows the scan is a few milliseconds; an
index on the same expression brings it back when that stops being true.

LIKE's own wildcards are escaped now — typing "100%" searched for
everything before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-13 21:07:22 +02:00
parent 272e8b1acf
commit 33d053ce75
3 changed files with 119 additions and 10 deletions

View File

@@ -7,6 +7,7 @@ import { Pagination } from "@/components/ui/Pagination";
import { StatusChip } from "@/components/ui/StatusChip";
import { currentUserId } from "@/lib/auth/session";
import { sql, withUser } from "@/lib/db";
import { istPersonalnummer, suchMuster } from "@/lib/employee-search";
import { derivedStatusFilter } from "@/lib/employee-status-filter";
import { fmtDate, todayIso } from "@/lib/format";
import { breadcrumbLabel, divisionOf, loadOrgMaps, subtreeOf, unitOf } from "@/lib/org";
@@ -87,7 +88,7 @@ export default async function EmployeesPage({ searchParams }: EmployeesPageProps
// `\d`, nicht `d`: der fehlende Backslash liess die Ziffernerkennung
// nie greifen — „1590" wurde als Name gesucht und fand nichts,
// während das Muster auf „ddd" ansprang.
if (/^\d+$/.test(term)) {
if (istPersonalnummer(term)) {
q = q.where("personnel_number", "=", Number(term));
} else {
// Wortweise statt am Stück, und **jedes** Wort muss irgendwo
@@ -109,17 +110,31 @@ export default async function EmployeesPage({ searchParams }: EmployeesPageProps
// Verglichen wird gegen den zusammengesetzten Namen, weil genau
// darauf der Trigramm-Index liegt (idx_employees_name_trgm).
// Getrennte Felder hätten ihn ungenutzt gelassen.
const woerter = term.split(/\s+/).filter(Boolean);
// Jedes Wort trifft am **Wortanfang**, nicht irgendwo mittendrin.
//
// Vorher wurde jedes Wort als Teilzeichenkette gesucht. Bei „Winkler
// H" traf das „H" auf T-h-omas, Kat-h-arina und CNC-Dre-h-er:in —
// die Suche gab alle sieben Winkler zurück, obwohl genau eine Hannah
// heisst. Je kürzer die Eingabe, desto unbrauchbarer wurde sie, und
// ein Anfangsbuchstabe ist die kürzeste sinnvolle Eingabe überhaupt.
//
// Gesucht wird über Vorname, Nachname und Position zusammen, wobei
// Trennzeichen als Wortgrenze gelten: „dreher" findet damit auch
// „CNC-Dreher:in". Ein Wort trifft, wenn der Heuhaufen damit beginnt
// oder ein Leerzeichen davorsteht.
//
// Der Trigramm-Index auf dem zusammengesetzten Namen greift hier
// nicht mehr — das ist der Preis. Bei knapp neunhundert Zeilen liest
// Postgres die Tabelle in wenigen Millisekunden; die Genauigkeit ist
// das wert, und bei Bedarf trägt ein Index auf demselben Ausdruck
// das später wieder.
const heuhaufen = sql<string>`translate(lower(first_name || ' ' || last_name || ' ' || job_title), '-/:.,', ' ')`;
q = q.where((eb) =>
eb.and(
woerter.map((wort) => {
// Als Parameter gebunden, nicht in die Abfrage geschrieben.
const like = `%${wort}%`;
return eb.or([
eb(sql<string>`first_name || ' ' || last_name`, "ilike", like),
eb("job_title", "ilike", like),
]);
})
// Als Parameter gebunden, nicht in die Abfrage geschrieben.
suchMuster(term).map(([amAnfang, nachLeerzeichen]) =>
eb.or([eb(heuhaufen, "like", amAnfang), eb(heuhaufen, "like", nachLeerzeichen)])
)
)
);
}