remove parquet output support
Nate has given up on getting wikiq to output parquet directly because it makes too many unpredictable large memory allocations. The workflow now outputs JSONL in a single pass and then uses spark (wikiq_spark) to index it as parquet in a second pass. Remove the parquet output path along with the machinery that existed only to support it: checkpoint files, resume temp-file merging, namespace partitioning, and file rotation (--partition-namespaces, --max-revisions-per-file). Resume support remains for JSONL output, where the resume point is derived from the last complete line of the output file. Also remove the parquet tests and baseline files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -42,17 +42,8 @@ class WikiqTester:
|
||||
else:
|
||||
shutil.rmtree(self.output)
|
||||
|
||||
# Also clean up resume-related files
|
||||
for temp_suffix in [".resume_temp", ".checkpoint", ".merged"]:
|
||||
temp_path = self.output + temp_suffix
|
||||
if os.path.exists(temp_path):
|
||||
if os.path.isfile(temp_path):
|
||||
os.remove(temp_path)
|
||||
else:
|
||||
shutil.rmtree(temp_path)
|
||||
|
||||
# For JSONL and Parquet, self.output is a file path. Create parent directory if needed.
|
||||
if out_format in ("jsonl", "parquet"):
|
||||
# For JSONL, self.output is a file path. Create parent directory if needed.
|
||||
if out_format == "jsonl":
|
||||
parent_dir = os.path.dirname(self.output)
|
||||
if parent_dir:
|
||||
os.makedirs(parent_dir, exist_ok=True)
|
||||
|
||||
Reference in New Issue
Block a user