Prepare wikiq for release: fixes, dependency cleanup, and packaging #2
99
README.md
99
README.md
@@ -1,46 +1,89 @@
|
||||
When you install this from git, you will need to first clone the
|
||||
repository:
|
||||
# mediawiki_dump_tools
|
||||
|
||||
git clone git://projects.mako.cc/mediawiki_dump_tools
|
||||
Tools for converting MediaWiki XML database dumps—the "history" dumps
|
||||
that include every revision of every page—into tabular datasets for
|
||||
research. The main tool is `wikiq`, a command line program that produces
|
||||
one row per revision with metadata such as the page, namespace,
|
||||
timestamp, editor, text size, and revert status, plus a range of
|
||||
optional computed columns.
|
||||
|
||||
From within the repository working directory, initiatlize and set up the
|
||||
submodule like:
|
||||
## Installation
|
||||
|
||||
git submodule init
|
||||
git submodule update
|
||||
wikiq requires Python 3.11 or later. Install it with pip from a clone of
|
||||
this repository:
|
||||
|
||||
Wikimedia dumps are usually in a compressed format such as 7z (most
|
||||
common), gz, or bz2. Wikiq uses your computer's compression software to
|
||||
read these files. Therefore wikiq depends on `7za`, `gzcat`, and `zcat`.
|
||||
git clone https://gitea.communitydata.science/collective/mediawiki_dump_tools.git
|
||||
cd mediawiki_dump_tools
|
||||
pip install .
|
||||
|
||||
# Dependencies
|
||||
[uv](https://docs.astral.sh/uv/) works too (`uv sync`) and is convenient
|
||||
for development, but it is not required.
|
||||
|
||||
These non-Python dependencies must be installed on your system for wikiq
|
||||
and its associated tests to work.
|
||||
One dependency, `pywikidiff2`, is installed from source and compiles a
|
||||
C++ extension, so a C++ compiler must be available at install time.
|
||||
|
||||
- 7zip
|
||||
- ffmpeg
|
||||
Wikimedia dumps are usually compressed as 7z (most common), gz, or bz2.
|
||||
wikiq reads these by running your system's decompression tools, so it
|
||||
depends on `7za`, `zcat`, and `bzcat` for those respective formats. On
|
||||
Debian or Ubuntu, `apt install 7zip` provides `7za`; the others are
|
||||
standard.
|
||||
|
||||
A new diff engine based on [wikidiff2](https://www.mediawiki.org/wiki/Wikidiff2)
|
||||
can be used for word-persistence. Wikiq can also output the diffs
|
||||
between each page revision. This requires installing Wikidiff 2 on your
|
||||
system. On Debian or Ubuntu Linux this can be done via.
|
||||
## Usage
|
||||
|
||||
`apt-get install php-wikidiff2`
|
||||
wikiq dump.xml.7z -o output/
|
||||
|
||||
You may have to also run. `sudo phpenmod wikidiff2`.
|
||||
Each output row describes one revision. The output format follows the
|
||||
`-o` argument: a path ending in `.jsonl` produces JSON Lines, anything
|
||||
else produces tab-separated values. With no dump file argument, wikiq
|
||||
reads XML on stdin and writes to stdout.
|
||||
|
||||
# Tests
|
||||
The most commonly useful options (`wikiq --help` describes them all):
|
||||
|
||||
To run tests:
|
||||
- `-n ID` limits output to a namespace, and can be given more than once.
|
||||
`-rr N` sets how many prior edits are checked when detecting reverts.
|
||||
- `--collapse-user` collapses each sequence of consecutive edits by the
|
||||
same user into a single row, which can address problems with text
|
||||
persistence measures.
|
||||
- `-p [METHOD]` computes content persistence measures for each revision:
|
||||
persistent token revisions, tokens added, and tokens removed. The
|
||||
available methods are `sequence` (the default), `segment` (robust to
|
||||
content moves but slower), `wikidiff2` (like segment, using the
|
||||
wikidiff2 diff engine), and `legacy` (the behavior of older research
|
||||
projects). Persistence is the slowest thing wikiq computes.
|
||||
- `-d` outputs a structured diff for each revision; `-t` outputs the
|
||||
full revision text.
|
||||
- `--external-links`, `--citations`, `--wikilinks`, `--templates`, and
|
||||
`--headings` parse each revision's wikitext and add a column with the
|
||||
extracted elements.
|
||||
- `-RP REGEX -RPl LABEL` searches revision text for a regular expression
|
||||
and reports the matches in a column named by the label; repeat the
|
||||
pair for multiple patterns. `-CP`/`-CPl` do the same for edit
|
||||
summaries. Adding `-RPc` or `-CPc` reports the number of matches
|
||||
instead of the matched text.
|
||||
- `--resume` continues an interrupted run from the last complete line of
|
||||
an existing JSONL output file. Combined with `--time-limit HOURS`,
|
||||
this supports processing very large dumps in bounded chunks.
|
||||
- `--print-schema` prints a Spark-compatible JSON schema for the
|
||||
configured output and exits.
|
||||
|
||||
python -m unittest test.Wikiq_Unit_Test
|
||||
## Tests
|
||||
|
||||
## TODO:
|
||||
From the repository root:
|
||||
|
||||
1. \[\] Output metadata about the run. What parameters were used? What
|
||||
versions of deltas?
|
||||
2. \[\] Url encoding by default
|
||||
uv run pytest test/
|
||||
|
||||
Run the suite from the root—some tests open fixture files by paths
|
||||
relative to it. The full suite processes several real dumps from
|
||||
`test/dumps/` and takes around fifteen minutes; expected outputs live in
|
||||
`test/baseline_output/`.
|
||||
|
||||
## Authors
|
||||
|
||||
wikiq is written and maintained by members of the [Community Data
|
||||
Science Collective](https://wiki.communitydata.science/). Contributors
|
||||
include Benjamin Mako Hill, Nathan TeBlunthuis, Will Beason, Sohyeon
|
||||
Hwang, and Kaylea Champion. It builds on earlier versions of the tool
|
||||
written in Python and in C++ by Benjamin Mako Hill.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
Reference in New Issue
Block a user