Machine learning models/Production/Spanish Wikiquote damaging edit

Model card
Model card
This page is an on-wiki machine learning model card.
	A model card is a document about a machine learning model that seeks to answer basic questions about the model.
Model Information Hub
Model creator(s)	Aaron Halfaker (User:EpochFail) and Amir Sarabadani
Model owner(s)	WMF Machine Learning Team (ml@wikimediafoundation.org)
Model interface	Ores homepage
Code	ORES Github, ORES training data, and ORES model binaries
Uses PII	No
In production?	Yes
Which projects?	Spanish Wikiquote
	This model uses data about a revision to predict the likelihood that the revision is damaging.
	v; t; e;

Motivation

Some goodfaith edits are damaging to an article, and not all damaging edits are in bad faith. This model (together with a goodfaith model) is intended to differentiate between edits that are intentionally harmful (badfaith/vandalism) and edits that are intended to be harmful (good edits/goodfaith damage).

This model helps to prioritize review of potentially damaging edits or vandalism. It provides a prediction on whether or not a given revision is damaging, and provides some probabilities to serve as a measure of its confidence level.

Users and uses

Use this model for

This model should be used for prioritizing the review and potential reversion of vandalism on Spanish Wikiquote.
This model should be used for detecting damaging contributions by editors on Spanish Wikiquote.

Don't use this model for

This model should not be used as an ultimate arbiter of whether or not an edit ought to be considered damaging.
The model should not be used outside of Spanish Wikiquote.

Current uses

Spanish Wikiquote uses the model as a service for facilitating efficient vandalism triage, edit reviews, or newcomer support.
On an individual basis, anyone can submit a properly-formatted API call to ORES for a given revision and get back the result of this model.

Example API call:

https://ores.wikimedia.org/v3/scores/eswikiquote/441558/damaging

Ethical considerations, caveats, and recommendations

Spanish Wikiquote decided to use this model. Over time, the model has been validated through use in the community.

This model is known to give newer editors higher probability of damaging edits.

Internal or external changes that could make this model deprecated or no longer usable are:

Data drift means training data for the model is no longer usable.
Doesn't meet desired performance metrics in production.
Spanish Wikiquote community decides to not use this model anymore.

Model

Performance

Test data confusion matrix:

Label	n	~True	~False
True	759	601	158
False	8163	611	7552

Test data sample rates:

Rate	Sample	Population
sample	0.085	0.915
population	0.087	0.913

Test data performance:

Statistic	True	False
match_rate	0.137	0.863
filter_rate	0.863	0.137
recall	0.792	0.925
precision	0.502	0.979
f1	0.615	0.951
accuracy	0.914	0.914
fpr	0.075	0.208
roc_auc	0.947	0.948
pr_auc	0.729	0.994

Implementation

Model architecture

{
    "type": "GradientBoosting",
    "params": {
        "scale": true,
        "center": true,
        "labels": [
            true,
            false
        ],
        "multilabel": false,
        "population_rates": null,
        "ccp_alpha": 0.0,
        "criterion": "friedman_mse",
        "init": null,
        "learning_rate": 0.01,
        "loss": "deviance",
        "max_depth": 7,
        "max_features": "log2",
        "max_leaf_nodes": null,
        "min_impurity_decrease": 0.0,
        "min_impurity_split": null,
        "min_samples_leaf": 1,
        "min_samples_split": 2,
        "min_weight_fraction_leaf": 0.0,
        "n_estimators": 700,
        "n_iter_no_change": null,
        "presort": "deprecated",
        "random_state": null,
        "subsample": 1.0,
        "tol": 0.0001,
        "validation_fraction": 0.1,
        "verbose": 0,
        "warm_start": false
    }
}

Output schema

{
    "title": "Scikit learn-based classifier score with probability",
    "type": "object",
    "properties": {
        "prediction": {
            "description": "The most likely label predicted by the estimator",
            "type": "boolean"
        },
        "probability": {
            "description": "A mapping of probabilities onto each of the potential output labels",
            "type": "object",
            "properties": {
                "true": {
                    "type": "number"
                },
                "false": {
                    "type": "number"
                }
            }
        }
    }
}

Example input and output

Input:

https://ores.wikimedia.org/v3/scores/eswikiquote/441558/damaging

Output:

{
    "eswikiquote": {
        "models": {
            "damaging": {
                "version": "0.5.0"
            }
        },
        "scores": {
            "441558": {
                "damaging": {
                    "score": {
                        "prediction": false,
                        "probability": {
                            "false": 0.980556466810999,
                            "true": 0.01944353318900106
                        }
                    }
                }
            }
        }
    }
}

Data

Data pipeline

Tabular data about edits is collected from the Mediawiki API, preprocessed (via log-transformations, joining with public editor data, etc.), and joined with user-generated goodfaith/damaging labels.

Training data

This model was trained using hand-labeled training data that is several years old.

Test data

The statistics reported here were calculated by selecting a random partition of the training data to hold out from the training process. The model then makes a prediction on that data, which is compared to the underlying ground truth.

Licenses

Code: MIT license
Model: MIT license

Citation

Cite this model card as:

@misc{
  Triedman_Bazira_2023_Spanish_Wikiquote_damaging,
  title={ Spanish Wikiquote damaging model card },
  author={ Triedman, Harold and Bazira, Kevin },
  year={ 2023 },
  url={ https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Spanish_Wikiquote_damaging_edit }
}