20
0

datasets: explicit partition counts and cache df between sorts

sort_and_write() now takes num_partitions (default from config or 200)
and passes it explicitly to repartition() instead of relying on
spark.sql.shuffle.partitions. comments_part2 defaults to 800 partitions
(~1GB/file matching the existing dataset density) and submissions_part2
to 300. Also cache the preprocessed DataFrame so the input is not
re-read and reshuffled twice when writing the by_subreddit and
by_author outputs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-05-29 14:48:05 -07:00
parent 9bc871285b
commit 172e66865c
3 changed files with 23 additions and 7 deletions

View File

@@ -15,7 +15,9 @@ from dumps_helper import SUBMISSIONS, sort_and_write
if __name__ == "__main__":
fire.Fire(lambda indir=None, out_by_subreddit=None, out_by_author=None:
fire.Fire(lambda indir=None, out_by_subreddit=None, out_by_author=None,
num_partitions=300:
sort_and_write(SUBMISSIONS, indir=indir,
out_by_subreddit=out_by_subreddit,
out_by_author=out_by_author))
out_by_author=out_by_author,
num_partitions=num_partitions))