Methods education

Understand the method. Use the evidence wisely.

Sense. Orient. Act. Learn. That is MP2’s operating model. This page supports Orient: it helps leaders ask what an analysis can answer, how it was tested, and where context and human judgment must enter.

Two questions, not one.

Purpose is the question. Method is the tool. Design is the reason to trust the answer: what was measured, when it was available, and which comparison makes the conclusion credible.

Layer 1 Purpose

Four jobs. Different evidence.

These purposes inform one another; they are not a mandatory sequence. Leaders often must act with incomplete evidence. The discipline is to state what is known, make the uncertainty visible, and learn from the action.

01

Descriptive

What happened?

Counts, rates, and trends in the records you already hold.

Gives you an accurate picture of what was recorded.
02

Diagnostic

What could explain it?

Testing explanations against timelines, local knowledge, and rival accounts.

Gives you evidence about mechanisms, with confidence tied to the design.
03

Predictive

What is likely next?

Using past patterns to estimate a probability for a period not yet seen.

Gives you decision support under uncertainty.
04

Causal

What did the action change?

Comparing an observed outcome with a credible estimate of what would have happened otherwise.

Gives you an effect estimate, conditional on the design and its assumptions.

Layer 2 Method

Five families sit underneath.

Not a ranking. Not stages. And not a one-to-one map onto the four purposes above.

  • Description and measurementRates, denominators, distributions, visualization, data quality.
  • Case study and qualitative inquiryInterviews, documents, process tracing, comparison across a few cases.
  • RegressionLinear, logistic, Poisson, and their relatives.
  • Machine learningTree ensembles and penalized models, judged on data they have not seen.
  • Research designs for causal estimationRandomized trials; difference-in-differences and other quasi-experiments.

Which method for which purpose?

Read each row as a guide to common uses, not a ban on other applications. The research design, outcome, and assumptions determine the claim. “Diagnose” here means investigate an organizational explanation, never diagnose a person.

Method families scored against four analytic purposes. Primary means the method is built for that job. Supports means it contributes but cannot carry the claim alone. With a design means it reaches that purpose only inside a research design that supplies a credible comparison. Not its job means another method is required.
Method family Describe Diagnose Predict Estimate an effect
Summary statistics & visualization Primary Supports Not its job Not its job
Case study & qualitative inquiry Supports Primary Not its job With a design
Regression (linear, logistic, Poisson) Supports Supports Primary With a design
Machine learning (tree ensembles, penalized models) Not its job Supports Primary With a design
Quasi-experimental designs (difference-in-differences, panel counterfactuals, and similar) Not its job Supports Not its job Primary
Randomized trial (the reference standard) Not its job Supports Not its job Primary

What moves a method right

A credible answer to compared with what? Random assignment can create comparable groups; attrition, spillovers, and implementation still matter. A quasi-experiment needs an explicit argument for its comparison. Statistical significance cannot supply a missing counterfactual.

What context adds

An alert can focus attention, but cannot name the cause. Review the timeline, organizational conditions, and reporting process. A rise in reports may reflect more harm, more willingness to report, or both. That distinction changes the appropriate response.

Where MP2 stops

MP2 describes, diagnoses, predicts, and evaluates effects. It does not automate the prescriptive step. A model narrows where leaders look; the decision stays with accountable people.

Open a method.

Seven modules, grouped by common use. Start with the visual and the leader question. Open the field notes for assumptions, execution, a worked example, failure modes, and primary-source reading. All module examples and code are synthetic; actual URFT evidence is labeled separately.

Family A

Describe and diagnose

Establish what was recorded and investigate how it came about. Description reveals the pattern; carefully designed case research can test explanations and mechanisms.

01 Pattern

Descriptive analysis

What pattern is visible in the records we have?

02 Explanation

Case study & qualitative inquiry

Why did this unfold the way it did, in this setting?

Family B

Predict

Estimate outcomes not yet observed at the decision point. Logistic regression models a binary outcome; Poisson regression models a count. XGBoost is a flexible model family that can handle either, depending on its objective. None supplies a causal explanation by itself.

03 Probability

Logistic regression

How likely is a yes-or-no outcome, given observed factors?

04 Counts

Poisson regression

How many recorded events might occur in a defined period?

05 Nonlinear pattern

XGBoost & tree ensembles

Do nonlinear thresholds or interactions improve prediction?

Family C

Estimate an effect

Pair a credible research design with an appropriate estimator. Difference-in-differences supplies a comparison strategy; panel estimators implement particular assumptions about the missing untreated path. Repeated observations alone do not establish causality.

06 Effect estimate

Difference-in-differences

Did change differ after an intervention relative to a comparison?

07 Counterfactual path

Panel-data estimators

What might treated units have looked like without the intervention?

From recent signals to an estimated probability.

One method, one purpose, made concrete. Logistic regression suits an outcome with two states: an incident was recorded in a unit-week, or it was not. The model combines observed factors into a score, then maps that score onto a probability between zero and one. Read what follows as prediction—the same equation placed inside a causal design would be a different claim entirely.

01
Define the unit and outcome

The public Unit Risk Forecasting Tool (URFT) paper analyzed company-, troop-, and battery-level weeks—not individual service members.

02
Weight observed signals

Recent recorded incidents and selected unit factors receive estimated coefficients.

z = β0 + β1x1 + … + βkxk
03
Map the score to probability

The logistic function bends the result into the zero-to-one range.

p = 1 / (1 + e−z)
04
Interpret with context

A probability can focus leader inquiry. It does not explain why conditions changed or prescribe an intervention.

Published logistic-regression plot of prior incidents and current-week incident probability A chart from the public URFT working paper. Jittered binary company-week observations appear around zero and one. A fitted blue curve rises as the number of serious incident reports in the prior three weeks increases. A gray band shows uncertainty.
Published Figure 4. The fitted curve describes an association between recent reports and a current-week recorded incident, not a causal effect or an individual forecast. The archived draft’s numerical probability examples require reconciliation with the specification behind this figure; they are not reproduced here. View the public paper

37company-sized units

5,439company-week observations

3 weeksrecent-incident window

Current weekreported prediction target

Accuracy is not the whole decision.

A model can find most recorded events and still produce more false positives than true positives. Read the error counts, the outcome’s base rate, and the quality of the probabilities together. This example uses published URFT counts; the teaching illustrations below are separate.

Reconstructed from Figure 8 of the public URFT working paper, at its stated 0.5 classification threshold. These counts describe that test set; they do not establish calibration, prospective validation, or performance in another formation. The threshold is not an operational alert rule.
93.0%
Overall accuracy

1,518 of 1,632 test weeks classified correctly. Most weeks recorded no incident, so correct negatives dominate the score.

95.7%
The always-negative baseline

Only 70 of 1,632 weeks recorded an incident. Always answering “no incident” achieves higher accuracy, but finds none of those 70 weeks. Accuracy alone cannot decide whether the model is useful.

91.4%
Recall

Of the recorded-incident weeks, 64 of 70 were classified positive.

37.2%
Precision

Of 172 positive classifications, 64 aligned with a recorded incident and 108 did not.

Read the probabilities and errors together

Precision and false-positive rate use different denominators. In the published matrix, 108 of 1,562 no-incident weeks were classified positive: a 6.9% false-positive rate. But 108 of 172 positive classifications were false positives: 62.8%. Both statements are true. A rare outcome can produce many false positives even when specificity is high.

Ranking and calibration answer different questions. Discrimination asks whether higher scores tend to identify outcome periods. Calibration asks whether estimated probabilities match observed frequencies across comparable periods. This confusion matrix, at one classification setting, cannot establish either the full ranking performance or probability calibration. See Van Calster et al. (2019).

A classification setting creates a workload. Changing it trades off misses and additional reviews. Compare models at a relevant common review burden, show uncertainty, and validate in the intended setting. These historical counts do not prescribe a threshold or a response.

Reporting is part of the outcome. “No recorded incident” is not proof that nothing happened. In a live support program, intervention may also change subsequent outcomes and reporting. An observed non-event after support does not by itself prove either a false alarm or a prevented event.

Try the arithmetic · synthetic illustration

Why rare outcomes change the picture.

How many positive classifications correspond to a recorded outcome? Hold a hypothetical classification rule and its assumed error rates fixed, then change how common the recorded outcome is.

Hypothetical observations
10,000 unit-periods
Sensitivity held fixed
80%
Specificity held fixed
90%
Classification rule
Unchanged

Sensitivity and specificity are held constant to isolate the base-rate effect. They can also change across settings and time. These are expected hypothetical counts, not URFT results.

Precision
29.6%
Hypothetical model accuracy
89.5%
Always-negative accuracy
95.0%
True positives among positive classificationsFalse positives
Expected classifications across 10,000 hypothetical unit-periods
Recorded outcomeClassified positiveClassified negative
Outcome recorded400True positives100False negatives
No outcome recorded950False positives8,550True negatives

At 5% prevalence, 400 of 1,350 positive classifications correspond to a recorded outcome: 29.6% precision. The other 950 are false positives relative to the recorded outcome.

See the calculation

Precision is true positives / (true positives + false positives). Let p be the recorded-outcome prevalence as a fraction. With the assumed sensitivity of 0.80 and specificity of 0.90:

Precision = (0.80 × p) / [(0.80 × p) + (0.10 × (1 − p))]

The number of hypothetical observations determines the counts. It cancels out of the precision calculation.

This illustrates predictive value, not calibration, causal impact, or the value of acting. It is not an operational threshold recommendation or an individual assessment. Review the metric definitions .

Seven questions before action.

A useful model is more than an accuracy score. It is a transparent chain from a decision need to an accountable human response.

  1. 01
    What decision will this inform?

    Name the question and the choice it could change. Ask the analyst to label the result: description, explanation, prediction, or effect estimate.

  2. 02
    What exactly is one observation?

    A company-week is an organizational observation, not an individual assessment. State the population, period, denominator, and limits on transfer to other formations.

  3. 03
    What entered the record—and what did not?

    Define the outcome, reporting delay, exclusions, and missingness. A confirmed zero and an absent report are different. More reporting can accompany better trust and access.

  4. 04
    What makes this a fair test?

    For prediction: untouched later data, useful baselines, and calibration. For an effect: a credible untreated comparison, clear timing, and uncertainty at the level of treatment assignment.

  5. 05
    What do the errors cost?

    Translate rates into expected reviews, misses, and resource demands. Include confidentiality, stigma, trust, and access to support—not only statistical scores.

  6. 06
    What finding would change our minds?

    Name the strongest rival explanation and the evidence that would challenge the result. Ask how conclusions change with plausible reporting, trend, and model assumptions.

  7. 07
    Who owns the response and the learning?

    Specify the human review, authorized supportive action, resource owner, and review date. Track implementation as well as outcomes, with clear conditions for revising or stopping the approach.

The essential boundary

A model can surface a signal. It cannot diagnose a person.

MP2 uses organizational evidence to focus inquiry, strengthen care, and improve systems. It does not support individual surveillance, automated personnel action, or clinical decisions.

Source status and the limits of this teaching page

The displayed Figure 4 is retained from the archived base-model draft. Its uncontrolled specification and the draft’s narrative probability example appear to differ; the exact calculation requires reconciliation. The page therefore teaches the qualitative association without repeating those numerical probabilities. This is a limitation of the archived material inspected, not a verified assessment of every version at the public DOI.

The original base model and later exploratory extensions use different datasets and targets. The later public account discusses a next-two-week binary model; that does not change the original paper’s current-week target. Historical impact estimates remain unresolved pending reconciliation and independent reproduction.

All module calculations, visual teaching diagrams, and code patterns are illustrative. They do not reproduce URFT analyses or propose operational thresholds. Public availability of an article does not authorize reuse of its underlying data, code, or sensitive operational detail.

MP2 method brief01 / 07

Education module · conceptual guide

Descriptive analysis

Turn records into a clear picture before trying to predict or explain.

Records summarized as a distribution and trendObserved records are grouped into comparable periods, shown as bars, and connected by a trend line. Variation remains visible around the summary.OBSERVED RATETIME PERIODPATTERN, NOT CAUSE
Conceptual illustration · not an MP2 result

Did the pattern change, or did the way we count it change?

Make the denominator part of the story.

Start with counts, rates, and variation across comparable units and periods. A useful description makes the record legible without pretending the record captures everything that happened.

It can tell us

What was recorded, where the pattern varies, and whether an apparent change survives a fair denominator and consistent definitions.

It cannot tell us

Why the pattern changed or whether a program caused it. A complete extract of administrative records can still miss unreported events.

Technical field notesMeasurement, worked example & source reading

01 Theory

Count, rate, and prevalence answer different questions.

A count describes volume. A rate relates events to a defined exposure, such as personnel-weeks. A proportion describes a share of an eligible population. Repeated reports are not necessarily different people; an event rate is not the percentage of personnel affected.

rate per 100 personnel-weeks = 100 × reports ÷ personnel-weeks
  • Define the record: one event, one report, or one person; these are not interchangeable.
  • Define exposure: who or what could contribute to the numerator, for how long.
  • Define coverage: a confirmed zero, a missing return, and an excluded unit need different codes.

02 Execute

Build the population before counting its events.

  1. Create a complete unit-period roster from an authorized exposure source, including periods with no events.
  2. Reconcile duplicates, event dates versus entry dates, changing categories, and unit reorganizations.
  3. Join event counts to that roster. Assign zero only where reporting completeness is confirmed; retain unknowns as missing.
  4. Show counts and rates together, then distributions, missingness, and sensitivity to reasonable time windows.
R / dplyrPython / pandasSQLPower BI

03 Illustrative code · R

# Synthetic schema; records already deduplicated.
# exposure: one row for EVERY eligible unit-week.
counts = dplyr::count(records, unit_id, week,
                      name = "reports")
unit_week = dplyr::left_join(
  exposure, counts, by = c("unit_id", "week"))
unit_week = dplyr::mutate(unit_week,
  reports = dplyr::if_else(reporting_complete,
    dplyr::coalesce(reports, 0L), NA_integer_),
  rate_100 = 100 * reports / person_weeks)

Require unique unit-week keys, positive exposure, and a verified completeness flag. No data are supplied or accessed. The example below demonstrates arithmetic, not an output from this snippet.

04 Worked example · synthetic

12 → 18reports rise; the rate stays at 1.2 per 100 personnel-weeks

In two fictional observation windows, 12 reports over 1,000 personnel-weeks become 18 over 1,500. The count rises 50%; exposure also rises 50%. Both rates are 1.2. The workload grew, but this comparison shows no rise in the recorded rate.

A rate is still a summary. Inspect variation across units and report uncertainty where the inference requires it; small denominators can produce unstable comparisons.

05 Failure mode

A cleaner chart can conceal a worse measure.

Excluding missing returns can make the data look more complete than they are. Changes in unit mix can move the overall average even when rates within each unit type do not change. Check both the aggregate and comparable groups, without exposing small cells or identifiable details.

MP2 guardrail: organizational summaries support inquiry and resources. They do not identify an individual at risk or rank commanders for punishment.

URFT Public descriptive foundation

One panel; two different denominators.

The public URFT working paper describes 37 company-sized units and 5,439 company-week observations. A company-week measures a unit over time, not personnel exposure. The 1,632-week test set shown on this page is a subset; its error rates must not be calculated using the full-panel denominator.

Defined recordsComplete unit-period rosterChecked summariesLeader inquiry

Read Primary sources