update README for the current tool
Describe what wikiq is and rewrite the installation, usage, and test instructions to match the current code: pip or uv installation from the gitea repository, the Python 3.11 requirement, output format selection by extension, and the current option set including persistence methods, wikitext extraction, regex matching and counting, and JSONL resume. Add an authors section crediting the Community Data Science Collective contributors and the earlier Python and C++ versions of the tool. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
99
README.md
99
README.md
@@ -1,46 +1,89 @@
|
|||||||
When you install this from git, you will need to first clone the
|
# mediawiki_dump_tools
|
||||||
repository:
|
|
||||||
|
|
||||||
git clone git://projects.mako.cc/mediawiki_dump_tools
|
Tools for converting MediaWiki XML database dumps—the "history" dumps
|
||||||
|
that include every revision of every page—into tabular datasets for
|
||||||
|
research. The main tool is `wikiq`, a command line program that produces
|
||||||
|
one row per revision with metadata such as the page, namespace,
|
||||||
|
timestamp, editor, text size, and revert status, plus a range of
|
||||||
|
optional computed columns.
|
||||||
|
|
||||||
From within the repository working directory, initiatlize and set up the
|
## Installation
|
||||||
submodule like:
|
|
||||||
|
|
||||||
git submodule init
|
wikiq requires Python 3.11 or later. Install it with pip from a clone of
|
||||||
git submodule update
|
this repository:
|
||||||
|
|
||||||
Wikimedia dumps are usually in a compressed format such as 7z (most
|
git clone https://gitea.communitydata.science/collective/mediawiki_dump_tools.git
|
||||||
common), gz, or bz2. Wikiq uses your computer's compression software to
|
cd mediawiki_dump_tools
|
||||||
read these files. Therefore wikiq depends on `7za`, `gzcat`, and `zcat`.
|
pip install .
|
||||||
|
|
||||||
# Dependencies
|
[uv](https://docs.astral.sh/uv/) works too (`uv sync`) and is convenient
|
||||||
|
for development, but it is not required.
|
||||||
|
|
||||||
These non-Python dependencies must be installed on your system for wikiq
|
One dependency, `pywikidiff2`, is installed from source and compiles a
|
||||||
and its associated tests to work.
|
C++ extension, so a C++ compiler must be available at install time.
|
||||||
|
|
||||||
- 7zip
|
Wikimedia dumps are usually compressed as 7z (most common), gz, or bz2.
|
||||||
- ffmpeg
|
wikiq reads these by running your system's decompression tools, so it
|
||||||
|
depends on `7za`, `zcat`, and `bzcat` for those respective formats. On
|
||||||
|
Debian or Ubuntu, `apt install 7zip` provides `7za`; the others are
|
||||||
|
standard.
|
||||||
|
|
||||||
A new diff engine based on [wikidiff2](https://www.mediawiki.org/wiki/Wikidiff2)
|
## Usage
|
||||||
can be used for word-persistence. Wikiq can also output the diffs
|
|
||||||
between each page revision. This requires installing Wikidiff 2 on your
|
|
||||||
system. On Debian or Ubuntu Linux this can be done via.
|
|
||||||
|
|
||||||
`apt-get install php-wikidiff2`
|
wikiq dump.xml.7z -o output/
|
||||||
|
|
||||||
You may have to also run. `sudo phpenmod wikidiff2`.
|
Each output row describes one revision. The output format follows the
|
||||||
|
`-o` argument: a path ending in `.jsonl` produces JSON Lines, anything
|
||||||
|
else produces tab-separated values. With no dump file argument, wikiq
|
||||||
|
reads XML on stdin and writes to stdout.
|
||||||
|
|
||||||
# Tests
|
The most commonly useful options (`wikiq --help` describes them all):
|
||||||
|
|
||||||
To run tests:
|
- `-n ID` limits output to a namespace, and can be given more than once.
|
||||||
|
`-rr N` sets how many prior edits are checked when detecting reverts.
|
||||||
|
- `--collapse-user` collapses each sequence of consecutive edits by the
|
||||||
|
same user into a single row, which can address problems with text
|
||||||
|
persistence measures.
|
||||||
|
- `-p [METHOD]` computes content persistence measures for each revision:
|
||||||
|
persistent token revisions, tokens added, and tokens removed. The
|
||||||
|
available methods are `sequence` (the default), `segment` (robust to
|
||||||
|
content moves but slower), `wikidiff2` (like segment, using the
|
||||||
|
wikidiff2 diff engine), and `legacy` (the behavior of older research
|
||||||
|
projects). Persistence is the slowest thing wikiq computes.
|
||||||
|
- `-d` outputs a structured diff for each revision; `-t` outputs the
|
||||||
|
full revision text.
|
||||||
|
- `--external-links`, `--citations`, `--wikilinks`, `--templates`, and
|
||||||
|
`--headings` parse each revision's wikitext and add a column with the
|
||||||
|
extracted elements.
|
||||||
|
- `-RP REGEX -RPl LABEL` searches revision text for a regular expression
|
||||||
|
and reports the matches in a column named by the label; repeat the
|
||||||
|
pair for multiple patterns. `-CP`/`-CPl` do the same for edit
|
||||||
|
summaries. Adding `-RPc` or `-CPc` reports the number of matches
|
||||||
|
instead of the matched text.
|
||||||
|
- `--resume` continues an interrupted run from the last complete line of
|
||||||
|
an existing JSONL output file. Combined with `--time-limit HOURS`,
|
||||||
|
this supports processing very large dumps in bounded chunks.
|
||||||
|
- `--print-schema` prints a Spark-compatible JSON schema for the
|
||||||
|
configured output and exits.
|
||||||
|
|
||||||
python -m unittest test.Wikiq_Unit_Test
|
## Tests
|
||||||
|
|
||||||
## TODO:
|
From the repository root:
|
||||||
|
|
||||||
1. \[\] Output metadata about the run. What parameters were used? What
|
uv run pytest test/
|
||||||
versions of deltas?
|
|
||||||
2. \[\] Url encoding by default
|
Run the suite from the root—some tests open fixture files by paths
|
||||||
|
relative to it. The full suite processes several real dumps from
|
||||||
|
`test/dumps/` and takes around fifteen minutes; expected outputs live in
|
||||||
|
`test/baseline_output/`.
|
||||||
|
|
||||||
|
## Authors
|
||||||
|
|
||||||
|
wikiq is written and maintained by members of the [Community Data
|
||||||
|
Science Collective](https://wiki.communitydata.science/). Contributors
|
||||||
|
include Benjamin Mako Hill, Nathan TeBlunthuis, Will Beason, Sohyeon
|
||||||
|
Hwang, and Kaylea Champion. It builds on earlier versions of the tool
|
||||||
|
written in Python and in C++ by Benjamin Mako Hill.
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user