diff --git a/README.md b/README.md index 54fbeed..720e92c 100644 --- a/README.md +++ b/README.md @@ -1,46 +1,89 @@ -When you install this from git, you will need to first clone the -repository: +# mediawiki_dump_tools - git clone git://projects.mako.cc/mediawiki_dump_tools +Tools for converting MediaWiki XML database dumps—the "history" dumps +that include every revision of every page—into tabular datasets for +research. The main tool is `wikiq`, a command line program that produces +one row per revision with metadata such as the page, namespace, +timestamp, editor, text size, and revert status, plus a range of +optional computed columns. -From within the repository working directory, initiatlize and set up the -submodule like: +## Installation - git submodule init - git submodule update +wikiq requires Python 3.11 or later. Install it with pip from a clone of +this repository: -Wikimedia dumps are usually in a compressed format such as 7z (most -common), gz, or bz2. Wikiq uses your computer's compression software to -read these files. Therefore wikiq depends on `7za`, `gzcat`, and `zcat`. + git clone https://gitea.communitydata.science/collective/mediawiki_dump_tools.git + cd mediawiki_dump_tools + pip install . -# Dependencies +[uv](https://docs.astral.sh/uv/) works too (`uv sync`) and is convenient +for development, but it is not required. -These non-Python dependencies must be installed on your system for wikiq -and its associated tests to work. +One dependency, `pywikidiff2`, is installed from source and compiles a +C++ extension, so a C++ compiler must be available at install time. -- 7zip -- ffmpeg +Wikimedia dumps are usually compressed as 7z (most common), gz, or bz2. +wikiq reads these by running your system's decompression tools, so it +depends on `7za`, `zcat`, and `bzcat` for those respective formats. On +Debian or Ubuntu, `apt install 7zip` provides `7za`; the others are +standard. -A new diff engine based on [wikidiff2](https://www.mediawiki.org/wiki/Wikidiff2) -can be used for word-persistence. Wikiq can also output the diffs -between each page revision. This requires installing Wikidiff 2 on your -system. On Debian or Ubuntu Linux this can be done via. +## Usage -`apt-get install php-wikidiff2` + wikiq dump.xml.7z -o output/ -You may have to also run. `sudo phpenmod wikidiff2`. +Each output row describes one revision. The output format follows the +`-o` argument: a path ending in `.jsonl` produces JSON Lines, anything +else produces tab-separated values. With no dump file argument, wikiq +reads XML on stdin and writes to stdout. -# Tests +The most commonly useful options (`wikiq --help` describes them all): -To run tests: +- `-n ID` limits output to a namespace, and can be given more than once. + `-rr N` sets how many prior edits are checked when detecting reverts. +- `--collapse-user` collapses each sequence of consecutive edits by the + same user into a single row, which can address problems with text + persistence measures. +- `-p [METHOD]` computes content persistence measures for each revision: + persistent token revisions, tokens added, and tokens removed. The + available methods are `sequence` (the default), `segment` (robust to + content moves but slower), `wikidiff2` (like segment, using the + wikidiff2 diff engine), and `legacy` (the behavior of older research + projects). Persistence is the slowest thing wikiq computes. +- `-d` outputs a structured diff for each revision; `-t` outputs the + full revision text. +- `--external-links`, `--citations`, `--wikilinks`, `--templates`, and + `--headings` parse each revision's wikitext and add a column with the + extracted elements. +- `-RP REGEX -RPl LABEL` searches revision text for a regular expression + and reports the matches in a column named by the label; repeat the + pair for multiple patterns. `-CP`/`-CPl` do the same for edit + summaries. Adding `-RPc` or `-CPc` reports the number of matches + instead of the matched text. +- `--resume` continues an interrupted run from the last complete line of + an existing JSONL output file. Combined with `--time-limit HOURS`, + this supports processing very large dumps in bounded chunks. +- `--print-schema` prints a Spark-compatible JSON schema for the + configured output and exits. - python -m unittest test.Wikiq_Unit_Test +## Tests -## TODO: +From the repository root: -1. \[\] Output metadata about the run. What parameters were used? What - versions of deltas? -2. \[\] Url encoding by default + uv run pytest test/ + +Run the suite from the root—some tests open fixture files by paths +relative to it. The full suite processes several real dumps from +`test/dumps/` and takes around fifteen minutes; expected outputs live in +`test/baseline_output/`. + +## Authors + +wikiq is written and maintained by members of the [Community Data +Science Collective](https://wiki.communitydata.science/). Contributors +include Benjamin Mako Hill, Nathan TeBlunthuis, Will Beason, Sohyeon +Hwang, and Kaylea Champion. It builds on earlier versions of the tool +written in Python and in C++ by Benjamin Mako Hill. ## License