isolation_forest_fire 5aad0a100366

Unsupervised anomaly detector for daily fire activity over the TerraSentinel study regions. Trained with IsolationForest on 27 strictly causal features β€” every rolling window ends at 1 preceding, so no feature can see the day it scores.

Provenance

Dataset commit e9fdff66d92411a5bdafc476d33db2369d42c227
Gold table gold.gold_fire_anomalies
Rows 1476
Regions greece_fire, iberia_fire
Date range 2024-09-28 β†’ 2026-10-05
Reference rule median/MAD z-score in gold.gold_fire_anomalies (is_anomaly, z >= 5)
Trained at 2026-10-05T12:56:56.484055+00:00
n_estimators 300
Contamination 0.025

The dataset commit is the point of this table: it is what lets you answer "which data produced this model".

Evaluation

No labels exist, so accuracy and F1 are not reported β€” they would be fabricated. What was measured:

Metric Value
Flagged days (contamination budget) 37.0000
Mean detections, flagged slice 2521.0000
Mean detections, rest 122.9402
Ratio (separation) 20.5059x
Agreement with the median/MAD rule 0.3243
Jaccard with the rule 0.1967
Score mass in the top slice 0.0452
Known-event regression 1.0000 (6/6 documented events)

Reference agreement. The comparator is gold_fire_anomalies.is_anomaly β€” the independent median/MAD rule (z β‰₯ 5) β€” not this model's own flags, so the numbers below can disagree and the model can be wrong. It flags 37 days, the rule flags 36, they share 12: precision 0.3243, reference recall 0.3333, Jaccard 0.1967.

Documented events used as a regression test

ml/validation/known_events.py holds real events with measured signatures. 3 of 6 were caught only by rank (below the serving threshold), which is reported rather than hidden:

  • iberia_fire 2025-08-15..2025-08-17, peak 13,329 β€” caught by the serving threshold
  • greece_fire 2025-08-12..2025-08-13, peak 2,005 β€” caught by rank only, below the serving threshold
  • iberia_fire 2026-02-24..2026-02-27, peak 1,184 β€” caught by rank only, below the serving threshold
  • iberia_fire 2026-07-03..2026-07-03, peak 2,238 β€” caught by the serving threshold
  • greece_fire 2024-09-30..2024-09-30, peak 808 β€” caught by rank only, below the serving threshold
  • iberia_fire 2025-07-26..2025-07-26, peak 172 β€” correctly not flagged (negative control)

Limitations

  • 2 regions, 2024-09-28 β†’ 2026-10-05. 1,476 training rows. This is a baseline, not a production dataset.
  • No labels. Every number above is distributional or agreement-based. Nothing here is a precision or recall against truth.
  • Instrument FRP is not comparable across sensors: MODIS mean FRP is ~100 MW where VIIRS is ~15 MW for the same fires. The model consumes frp_sum and frp_per_detection pooled across instruments, so intensity features carry an instrument-mix confound.
  • The percentile score is relative to the training distribution. A genuinely new regime (a year far outside the training range) will saturate the percentile.

Usage

from ml.bundle import load_bundle, score_to_percentile

bundle = load_bundle("path/or/hub/snapshot")
raw = -bundle.model.score_samples(X[bundle.feature_columns])
percentiles, comparable = score_to_percentile(bundle, raw)

Anomaly is percentile >= 0.9750.

Attribution

Fire detections: NASA FIRMS (MODIS and VIIRS active fire products). Sea-surface temperature precursor: NOAA OISST v2.1. Imagery (not used by this model, but by the pipeline): Copernicus Sentinel.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support