Benjamin Mako Hill 5fb7b5596c remove orphaned regextest.tsv baseline
No test reads this file; the regex tests use the basic_regextest and
capturegroup_regextest baselines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 16:23:52 -07:00
2026-08-06 16:17:14 -07:00
2026-08-06 16:17:14 -07:00
2026-08-06 16:23:52 -07:00

When you install this from git, you will need to first clone the repository:

git clone git://projects.mako.cc/mediawiki_dump_tools

From within the repository working directory, initiatlize and set up the submodule like:

git submodule init
git submodule update

Wikimedia dumps are usually in a compressed format such as 7z (most common), gz, or bz2. Wikiq uses your computer's compression software to read these files. Therefore wikiq depends on 7za, gzcat, and zcat.

Dependencies

These non-Python dependencies must be installed on your system for wikiq and its associated tests to work.

  • 7zip
  • ffmpeg

A new diff engine based on wikidiff2 can be used for word-persistence. Wikiq can also output the diffs between each page revision. This requires installing Wikidiff 2 on your system. On Debian or Ubuntu Linux this can be done via.

apt-get install php-wikidiff2

You may have to also run. sudo phpenmod wikidiff2.

Tests

To run tests:

python -m unittest test.Wikiq_Unit_Test

TODO:

  1. [] Output metadata about the run. What parameters were used? What versions of deltas?
  2. [] Url encoding by default
Description
command line tool to convert MediaWiki XML database dumps into tabular datasets for research
Readme 107 MiB
Languages
Python 100%