Descriptive
What happened?
Counts, rates, and trends in the records you already hold.
Methods education
Sense. Orient. Act. Learn. That is MP2’s operating model. This page supports Orient: it helps leaders ask what an analysis can answer, how it was tested, and where context and human judgment must enter.
01 Purpose and method
Purpose is the question. Method is the tool. Design is the reason to trust the answer: what was measured, when it was available, and which comparison makes the conclusion credible.
Layer 1 Purpose
These purposes inform one another; they are not a mandatory sequence. Leaders often must act with incomplete evidence. The discipline is to state what is known, make the uncertainty visible, and learn from the action.
What happened?
Counts, rates, and trends in the records you already hold.
What could explain it?
Testing explanations against timelines, local knowledge, and rival accounts.
What is likely next?
Using past patterns to estimate a probability for a period not yet seen.
What did the action change?
Comparing an observed outcome with a credible estimate of what would have happened otherwise.
Layer 2 Method
Not a ranking. Not stages. And not a one-to-one map onto the four purposes above.
02 The crosswalk
Read each row as a guide to common uses, not a ban on other applications. The research design, outcome, and assumptions determine the claim. “Diagnose” here means investigate an organizational explanation, never diagnose a person.
| Method family | Describe | Diagnose | Predict | Estimate an effect |
|---|---|---|---|---|
| Summary statistics & visualization | Primary | Supports | Not its job | Not its job |
| Case study & qualitative inquiry | Supports | Primary | Not its job | With a design |
| Regression (linear, logistic, Poisson) | Supports | Supports | Primary | With a design |
| Machine learning (tree ensembles, penalized models) | Not its job | Supports | Primary | With a design |
| Quasi-experimental designs (difference-in-differences, panel counterfactuals, and similar) | Not its job | Supports | Not its job | Primary |
| Randomized trial (the reference standard) | Not its job | Supports | Not its job | Primary |
What moves a method right
A credible answer to compared with what? Random assignment can create comparable groups; attrition, spillovers, and implementation still matter. A quasi-experiment needs an explicit argument for its comparison. Statistical significance cannot supply a missing counterfactual.
What context adds
An alert can focus attention, but cannot name the cause. Review the timeline, organizational conditions, and reporting process. A rise in reports may reflect more harm, more willingness to report, or both. That distinction changes the appropriate response.
Where MP2 stops
MP2 describes, diagnoses, predicts, and evaluates effects. It does not automate the prescriptive step. A model narrows where leaders look; the decision stays with accountable people.
03 Method modules
Seven modules, grouped by common use. Start with the visual and the leader question. Open the field notes for assumptions, execution, a worked example, failure modes, and primary-source reading. All module examples and code are synthetic; actual URFT evidence is labeled separately.
Family A
Establish what was recorded and investigate how it came about. Description reveals the pattern; carefully designed case research can test explanations and mechanisms.
01 Pattern
02 Explanation
Family B
Estimate outcomes not yet observed at the decision point. Logistic regression models a binary outcome; Poisson regression models a count. XGBoost is a flexible model family that can handle either, depending on its objective. None supplies a causal explanation by itself.
03 Probability
04 Counts
05 Nonlinear pattern
Family C
Pair a credible research design with an appropriate estimator. Difference-in-differences supplies a comparison strategy; panel estimators implement particular assumptions about the missing untreated path. Repeated observations alone do not establish causality.
06 Effect estimate
07 Counterfactual path
04 Worked example · Logistic regression, used predictively
One method, one purpose, made concrete. Logistic regression suits an outcome with two states: an incident was recorded in a unit-week, or it was not. The model combines observed factors into a score, then maps that score onto a probability between zero and one. Read what follows as prediction—the same equation placed inside a causal design would be a different claim entirely.
The public Unit Risk Forecasting Tool (URFT) paper analyzed company-, troop-, and battery-level weeks—not individual service members.
Recent recorded incidents and selected unit factors receive estimated coefficients.
z = β0 + β1x1 + … + βkxkThe logistic function bends the result into the zero-to-one range.
p = 1 / (1 + e−z)A probability can focus leader inquiry. It does not explain why conditions changed or prescribe an intervention.
37company-sized units
5,439company-week observations
3 weeksrecent-incident window
Current weekreported prediction target
05 Performance and error
A model can find most recorded events and still produce more false positives than true positives. Read the error counts, the outcome’s base rate, and the quality of the probabilities together. This example uses published URFT counts; the teaching illustrations below are separate.
Rows: recorded outcome
Columns: model classification
No incident
Incident
No incident
Incident
1,518 of 1,632 test weeks classified correctly. Most weeks recorded no incident, so correct negatives dominate the score.
Only 70 of 1,632 weeks recorded an incident. Always answering “no incident” achieves higher accuracy, but finds none of those 70 weeks. Accuracy alone cannot decide whether the model is useful.
Of the recorded-incident weeks, 64 of 70 were classified positive.
Of 172 positive classifications, 64 aligned with a recorded incident and 108 did not.
Precision and false-positive rate use different denominators. In the published matrix, 108 of 1,562 no-incident weeks were classified positive: a 6.9% false-positive rate. But 108 of 172 positive classifications were false positives: 62.8%. Both statements are true. A rare outcome can produce many false positives even when specificity is high.
Ranking and calibration answer different questions. Discrimination asks whether higher scores tend to identify outcome periods. Calibration asks whether estimated probabilities match observed frequencies across comparable periods. This confusion matrix, at one classification setting, cannot establish either the full ranking performance or probability calibration. See Van Calster et al. (2019).
A classification setting creates a workload. Changing it trades off misses and additional reviews. Compare models at a relevant common review burden, show uncertainty, and validate in the intended setting. These historical counts do not prescribe a threshold or a response.
Reporting is part of the outcome. “No recorded incident” is not proof that nothing happened. In a live support program, intervention may also change subsequent outcomes and reporting. An observed non-event after support does not by itself prove either a false alarm or a prevented event.
How many positive classifications correspond to a recorded outcome? Hold a hypothetical classification rule and its assumed error rates fixed, then change how common the recorded outcome is.
Sensitivity and specificity are held constant to isolate the base-rate effect. They can also change across settings and time. These are expected hypothetical counts, not URFT results.
| Recorded outcome | Classified positive | Classified negative |
|---|---|---|
| Outcome recorded | 400True positives | 100False negatives |
| No outcome recorded | 950False positives | 8,550True negatives |
At 5% prevalence, 400 of 1,350 positive classifications correspond to a recorded outcome: 29.6% precision. The other 950 are false positives relative to the recorded outcome.
Precision is true positives / (true positives + false positives). Let p be the recorded-outcome prevalence as a fraction. With the assumed sensitivity of 0.80 and specificity of 0.90:
Precision = (0.80 × p) / [(0.80 × p) + (0.10 × (1 − p))]
The number of hypothetical observations determines the counts. It cancels out of the precision calculation.
This illustrates predictive value, not calibration, causal impact, or the value of acting. It is not an operational threshold recommendation or an individual assessment. Review the metric definitions .
06 Read a model
A useful model is more than an accuracy score. It is a transparent chain from a decision need to an accountable human response.
Name the question and the choice it could change. Ask the analyst to label the result: description, explanation, prediction, or effect estimate.
A company-week is an organizational observation, not an individual assessment. State the population, period, denominator, and limits on transfer to other formations.
Define the outcome, reporting delay, exclusions, and missingness. A confirmed zero and an absent report are different. More reporting can accompany better trust and access.
For prediction: untouched later data, useful baselines, and calibration. For an effect: a credible untreated comparison, clear timing, and uncertainty at the level of treatment assignment.
Translate rates into expected reviews, misses, and resource demands. Include confidentiality, stigma, trust, and access to support—not only statistical scores.
Name the strongest rival explanation and the evidence that would challenge the result. Ask how conclusions change with plausible reporting, trend, and model assumptions.
Specify the human review, authorized supportive action, resource owner, and review date. Track implementation as well as outcomes, with clear conditions for revising or stopping the approach.
The essential boundary
MP2 uses organizational evidence to focus inquiry, strengthen care, and improve systems. It does not support individual surveillance, automated personnel action, or clinical decisions.
07 Continue your education
Start with the public application, then use each module’s primary-source reading to examine the method. A source’s publication status and the strength of its design are separate questions.
The displayed Figure 4 is retained from the archived base-model draft. Its uncontrolled specification and the draft’s narrative probability example appear to differ; the exact calculation requires reconciliation. The page therefore teaches the qualitative association without repeating those numerical probabilities. This is a limitation of the archived material inspected, not a verified assessment of every version at the public DOI.
The original base model and later exploratory extensions use different datasets and targets. The later public account discusses a next-two-week binary model; that does not change the original paper’s current-week target. Historical impact estimates remain unresolved pending reconciliation and independent reproduction.
All module calculations, visual teaching diagrams, and code patterns are illustrative. They do not reproduce URFT analyses or propose operational thresholds. Public availability of an article does not authorize reuse of its underlying data, code, or sensitive operational detail.