More records about the same source story
Group exact source URL + encoded occurrence date. Record count stays visible, but it is not multiplied into review priority. Ten coded mentions are not ten closures.
MODEL NOTES / SCREENING 2.0
Blackdeck separates source observations, inference and unknown quantities. Every score is a review aid, not a claim of economic loss.
A GDELT row with an action, actors and location. Many rows can come from the same news article.
These records are grouped for display. A story can contain several actions and therefore several taxonomy labels.
Same URL does not guarantee one real incident; different URLs may describe the same incident. Cross-publisher semantic deduplication is not yet complete.
Review priority = S × E × N × R
max(0, −Goldstein) × 10, capped at 100. An event-type reference value, not observed magnitude. A small and a large protest can have the same code value.0 without an eligible pathway; 0.30 for URL-only mechanism + asset clues; 0.45 for a coded mechanism + URL asset clue; 0.65 for the specific coded economic-restriction pathway. These are design weights, not calibrated probabilities.2^(−age_hours / 72), relative to the source snapshot cutoff. Old source data is dated visibly; the website does not present it as a live measurement.We removed the old “Global Risk Index” from the landing page. Averaging coded severity is not a defensible measurement of worldwide business loss. Likewise, correlation between a score and the mentions/source counts already inside that score is not an independent impact study.
REPETITION MODEL
Group exact source URL + encoded occurrence date. Record count stays visible, but it is not multiplied into review priority. Ten coded mentions are not ten closures.
Compare the country, actor pair and set of root event types with earlier story groups. The comparison uses earlier reporting dates within seven days. It is a similarity heuristic, not proof that the actual incident is identical.
N₀ = max(0.20, 1 / [1 + 0.5 × ln(1 + n)])
Here, n is the number of similar prior story groups. For an already seen source URL, the base new-information weight is capped at 0.25.
Illustrative formula outputs, not measured reactions from markets or supply chains.
Restore attention if the coded type severity exceeds the earlier maximum, or a newly matched asset or transmission mechanism appears.
N = N₀ + (1 − N₀) × material_change
material_change is the largest of: severity increase / 25 (capped at 1), 0.70 for a new asset clue, and 0.50 for a new mechanism clue. A large coded escalation can restore the full novelty weight. Evidence still requires review.
Blackdeck also measures how persistently a similar subject appears across reporting dates in the last 14 days. Repeated records on the same date count only once in this persistence measure.
A = Σ 2^(−days_ago / 7)
Reporting persistence = 100 × (1 − exp[−A / 5])
This is reporting persistence, not outage duration. Several days of reporting cannot establish that a port remained closed throughout those days. The value is not used as a fabricated cumulative loss.
A port strike may carry protest, labor, access and shipping labels. It still counts as one story group, and each impact edge keeps its own supporting evidence.
No estimated production loss or price return is inserted just to complete a dashboard. Unmapped events may still be important.
The map uses machine-coded geographic centroids. Captions derived from URLs are explicitly identified. Neither proves the location of an affected facility.
The novelty and evidence parameters are a transparent initial design. They are not trained or validated business-impact coefficients. Future calibration requires labeled incidents and observed outcomes.
GDELT: Goldstein describes event types, not the scale of a particular event ↗
Tetlock (2011): investor responses to stale news ↗
The study concerns financial-news responses. It does not validate these coefficients, or imply that repeated conflict always causes less physical damage. The present formulas are Blackdeck design choices.