Benjamin Mako Hill 1125194093 add the project metadata a published package needs
The package declared no authors, no URLs, no classifiers and no
keywords, so a PyPI page for it would be blank and it would not turn up
in a search. Fill those in.

Contributors are listed by name only. They are the same people the
README credits, so this discloses nothing new, and publishing other
people's email addresses to PyPI is not ours to decide.

Also mirror the dev dependencies into [project.optional-dependencies].
[dependency-groups] is PEP 735, which only uv reads, so there was no way
for a pip user to install what the tests need; `pip install -e '.[dev]'`
now works. The two lists have to be kept in step by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 7b3c0f9a80e39990df90124a7c592a80d93e0d91)
2026-08-13 16:53:52 -07:00
2026-08-13 13:35:14 -07:00

mediawiki_dump_tools

Tools for converting MediaWiki XML database dumps—the "history" dumps that include every revision of every page—into tabular datasets for research. The main tool is wikiq, a command line program that produces one row per revision with metadata such as the page, namespace, timestamp, editor, text size, and revert status, plus a range of optional computed columns.

Installation

wikiq requires Python 3.11 or later. Install it with pip from a clone of this repository:

git clone https://gitea.communitydata.science/collective/mediawiki_dump_tools.git
cd mediawiki_dump_tools
pip install .

uv works too (uv sync) and is convenient for development, but it is not required.

One dependency, pywikidiff2, is installed from source and compiles a C++ extension, so a C++ compiler must be available at install time.

Wikimedia dumps are usually compressed as 7z (most common), gz, or bz2. wikiq reads these by running your system's decompression tools, so it depends on 7za, zcat, and bzcat for those respective formats. On Debian or Ubuntu, apt install 7zip provides 7za; the others are standard.

Usage

wikiq dump.xml.7z -o output/

Each output row describes one revision. The output format follows the -o argument: a path ending in .jsonl produces JSON Lines, anything else produces tab-separated values. With no dump file argument, wikiq reads XML on stdin and writes to stdout.

The title and namespace columns are page-level identity values as of the time of export: a moved page's entire history carries its final title, and namespace derives from that title. Redirect status, by contrast, is determined per revision from each revision's own text: the revision_is_redirect and revision_redirect_target columns record whether a revision's text begins with a redirect directive and where it points. #REDIRECT is recognized on every wiki; wikis that also use a localized keyword can add it with --redirect-aliases (for example, --redirect-aliases OMDIRIGERING for Swedish wikis).

The most commonly useful options (wikiq --help describes them all):

  • -n ID limits output to a namespace, and can be given more than once. -rr N sets how many prior edits are checked when detecting reverts.
  • --collapse-user collapses each sequence of consecutive edits by the same user into a single row, which can address problems with text persistence measures.
  • -p [METHOD] computes content persistence measures for each revision: persistent token revisions, tokens added, and tokens removed. The available methods are sequence (the default), segment (robust to content moves but slower), wikidiff2 (like segment, using the wikidiff2 diff engine), and legacy (the behavior of older research projects). Persistence is the slowest thing wikiq computes.
  • -d outputs a structured diff for each revision; -t outputs the full revision text.
  • --external-links, --citations, --wikilinks, --templates, and --headings parse each revision's wikitext and add a column with the extracted elements.
  • -RP REGEX -RPl LABEL searches revision text for a regular expression and reports the matches in a column named by the label; repeat the pair for multiple patterns. -CP/-CPl do the same for edit summaries. Adding -RPc or -CPc reports the number of matches instead of the matched text.
  • --resume continues an interrupted run from the last complete line of an existing JSONL output file. Combined with --time-limit HOURS, this supports processing very large dumps in bounded chunks.
  • --print-schema prints a Spark-compatible JSON schema for the configured output and exits.

Tests

From the repository root:

uv run pytest test/

Run the suite from the root—some tests open fixture files by paths relative to it. The full suite processes several real dumps from test/dumps/ and takes around fifteen minutes; expected outputs live in test/baseline_output/.

Authors

wikiq is written and maintained by members of the Community Data Science Collective. Contributors include Benjamin Mako Hill, Nathan TeBlunthuis, Will Beason, Sohyeon Hwang, and Kaylea Champion. It builds on earlier versions of the tool written in Python and in C++ by Benjamin Mako Hill.

License

wikiq is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version. See COPYING for the full text.

Description
command line tool to convert MediaWiki XML database dumps into tabular datasets for research
Readme 107 MiB
Languages
Python 100%