Nate has given up on getting wikiq to output parquet directly because it makes too many unpredictable large memory allocations. The workflow now outputs JSONL in a single pass and then uses spark (wikiq_spark) to index it as parquet in a second pass. Remove the parquet output path along with the machinery that existed only to support it: checkpoint files, resume temp-file merging, namespace partitioning, and file rotation (--partition-namespaces, --max-revisions-per-file). Resume support remains for JSONL output, where the resume point is derived from the last complete line of the output file. Also remove the parquet tests and baseline files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
29 KiB
29 KiB