Benjamin Mako Hill 77f185888f run the whole suite from runtest.sh
It invoked pytest on test_wiki_diff_matcher.py alone, so it exercised 19
of the 69 tests despite its name, and it required uv, which contradicts
the README's statement that uv is optional.

Point it at test/ and use whichever python is on the path. It cds to the
repository root first, since several tests open fixture files by paths
relative to it, and passes any extra arguments through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 7fff65c23d7c7a2868abccb235ec376fe093d4c1)
2026-08-13 16:53:52 -07:00
2026-08-13 13:35:14 -07:00
2026-08-13 16:53:52 -07:00
2026-08-13 16:53:52 -07:00

mediawiki_dump_tools

Tools for converting MediaWiki XML database dumps—the "history" dumps that include every revision of every page—into tabular datasets for research. The main tool is wikiq, a command line program that produces one row per revision with metadata such as the page, namespace, timestamp, editor, text size, and revert status, plus a range of optional computed columns.

Installation

wikiq requires Python 3.11 or later. Install it with pip from a clone of this repository:

git clone https://gitea.communitydata.science/collective/mediawiki_dump_tools.git
cd mediawiki_dump_tools
pip install .

uv works too (uv sync) and is convenient for development, but it is not required.

This installs everything needed for the columns most people use, and requires no compiler. Two features carry heavier dependencies and are therefore optional.

--diff and -p wikidiff2 need pywikidiff2, a Python binding for MediaWiki's wikidiff2 diff engine. It is not on PyPI and compiles a C++ extension, so it needs a C++ compiler and libthai:

pip install 'pywikidiff2 @ git+https://gitea.communitydata.science/groceryheist/pywikidiff2.git'

-p legacy needs mediawiki-utilities, which is only useful for reproducing results from older research projects:

pip install '.[legacy]'

wikiq tells you which of these to install if you use an option that needs one. The other persistence methods, including the default -p sequence, need neither.

Wikimedia dumps are usually compressed as 7z (most common), gz, or bz2. wikiq reads these by running your system's decompression tools, so it depends on 7za, zcat, and bzcat for those respective formats. On Debian or Ubuntu, apt install 7zip provides 7za; the others are standard.

Usage

wikiq dump.xml.7z -o output/

Each output row describes one revision. The output format follows the -o argument: a path ending in .jsonl produces JSON Lines, anything else produces tab-separated values. With no dump file argument, wikiq reads XML on stdin and writes to stdout.

The title and namespace columns are page-level identity values as of the time of export: a moved page's entire history carries its final title, and namespace derives from that title. Redirect status, by contrast, is determined per revision from each revision's own text: the revision_is_redirect and revision_redirect_target columns record whether a revision's text begins with a redirect directive and where it points. #REDIRECT is recognized on every wiki; wikis that also use a localized keyword can add it with --redirect-aliases (for example, --redirect-aliases OMDIRIGERING for Swedish wikis).

The most commonly useful options (wikiq --help describes them all):

  • -n ID limits output to a namespace, and can be given more than once. -rr N sets how many prior edits are checked when detecting reverts.
  • --collapse-user collapses each sequence of consecutive edits by the same user into a single row, which can address problems with text persistence measures.
  • -p [METHOD] computes content persistence measures for each revision: persistent token revisions, tokens added, and tokens removed. The available methods are sequence (the default), segment (robust to content moves but slower), wikidiff2 (like segment, using the wikidiff2 diff engine), and legacy (the behavior of older research projects). Persistence is the slowest thing wikiq computes.
  • -d outputs a structured diff for each revision; -t outputs the full revision text.
  • --external-links, --citations, --wikilinks, --templates, and --headings parse each revision's wikitext and add a column with the extracted elements.
  • -RP REGEX -RPl LABEL searches revision text for a regular expression and reports the matches in a column named by the label; repeat the pair for multiple patterns. -CP/-CPl do the same for edit summaries.
  • --resume continues an interrupted run from the last complete line of an existing JSONL output file. Combined with --time-limit HOURS, this supports processing very large dumps in bounded chunks.
  • --print-schema prints a Spark-compatible JSON schema for the configured output and exits.

Tests

From the repository root:

pip install pytest pandas pytest-asyncio pytest-benchmark
pytest test/

Run the suite from the root—some tests open fixture files by paths relative to it. The full suite processes several real dumps from test/dumps/; expected outputs live in test/baseline_output/.

Tests covering --diff, -p wikidiff2, and -p legacy skip when their optional dependency is absent, so a base install reports skips rather than failures.

Authors

wikiq is written and maintained by members of the Community Data Science Collective. Contributors include Benjamin Mako Hill, Nathan TeBlunthuis, Will Beason, Sohyeon Hwang, and Kaylea Champion. It builds on earlier versions of the tool written in Python and in C++ by Benjamin Mako Hill.

License

wikiq is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version. See COPYING for the full text.

Description
command line tool to convert MediaWiki XML database dumps into tabular datasets for research
Readme 107 MiB
Languages
Python 100%