fix(articles, mana-ai): rollout-block hardening for sync_changes projections

Four cross-cutting fixes that make the bulk-import worker safe to run
under real production load. All four were called out as live-rollout
risks in the post-ship review of docs/plans/articles-bulk-import.md.

#1 — Same fieldMetaTime bug fixed in mana-ai
   The articles fix in 054b9e5be hoists the helper to its own file
   `apps/api/src/modules/articles/field-meta.ts`. The same naive
   `rowFM[k] >= localTime` LWW comparison existed in three more
   projections under services/mana-ai (missions-projection,
   snapshot-refresh, agents-projection). Once any F3 stamp lands
   beside a legacy-string stamp, the comparison evaluates
   `'[object Object]' >= 'ISO-…'` (false) and the older value wins.
   New `services/mana-ai/src/db/field-meta.ts` — same helper,
   deliberately duplicated (each service treats sync_changes as a
   read-only event log; sharing infra across services is out of
   scope here). All 61 mana-ai bun tests still pass.

#2 — Stale 'extracting' items recycle
   If the worker dies mid-fetch (OOM, pod restart), items stay in
   state='extracting' forever and the job never completes. New sweep
   at the start of `processOneJob`: items whose lastAttemptAt is
   older than 5 minutes get bounced back to 'pending' so the next
   tick re-claims them. STALE_EXTRACTING_MS tuned for the 15s
   shared-rss fetch + JSDOM-parse worst case.

#3 — Pickup-row GC
   Every 30 ticks (~once per minute) the worker hard-deletes
   articleExtractPickup rows older than 24h. Without this a stuck
   pickup-consumer (all tabs closed, Web-Lock mismatch) would let
   sync_changes accumulate without bound. Logs the row count when
   non-zero so we can spot stuck consumers in the wild.

#4 — DRY consent-wall heuristic
   Identical CONSENT_KEYWORDS + threshold lived in routes.ts AND
   import-extractor.ts. Hoisted to
   `apps/api/src/modules/articles/consent-wall.ts`; both call sites
   now share one heuristic.

Plan: docs/plans/articles-bulk-import.md.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Till JS 2026-04-29 00:53:39 +02:00
parent e99fea1938
commit b297f68ee4
10 changed files with 223 additions and 69 deletions

View file

@ -19,34 +19,20 @@
*/
import { getSyncConnection } from '../../mcp/sync-db';
import { fieldMetaTime } from './field-meta';
type Row = Record<string, unknown>;
/**
* `field_meta` is one of two shapes on the wire:
* - Legacy plaintext writes: `{[fieldName]: ISOString}`
* - Field-meta-overhaul writes: `{[fieldName]: {at, actor, origin}}`
* `fieldMetaTime()` below normalises both into the comparable ISO string.
*/
interface ChangeRow {
user_id: string;
record_id: string;
op: string;
data: Row | null;
/** See `./field-meta.ts` wire shape is two-tone (legacy ISO string
* vs. F3 `{at, actor, origin}` object). */
field_meta: Record<string, unknown> | null;
created_at: Date;
}
/** Pull the timestamp out of either shape. Falls back to empty string
* so the LWW comparison never throws on undefined. */
function fieldMetaTime(meta: unknown): string {
if (typeof meta === 'string') return meta;
if (meta && typeof meta === 'object') {
const at = (meta as { at?: unknown }).at;
if (typeof at === 'string') return at;
}
return '';
}
export interface ImportJobRow {
id: string;
userId: string;