How we measure framing
Every article on this site is scored by a language model. Models can be biased too, so the only useful answer to “how do you know it isn’t?” is a measurement. This page is that measurement, and it is updated from the same database that powers the scores.
Symmetry error
Not measured yet
Average score movement when we swap the political sides. Lower is better; 0 is perfect.
Direction flips
Not measured yet
How often the score failed to mirror after the swap. Lower is better.
Human agreement
Not measured yet
How often our score lands on the same side as a human labeller. Higher is better.
The scorecard
Five measurements of how well this works, each with the number of things it was measured over and the date it was last taken. Where a measurement has not been made, the row says so in words rather than showing a zero — an unmeasured cell that reads as a score is the one failure that would make this page worse than not having it.
| What we measured | Result | Sample | Last measured |
|---|---|---|---|
Agreement with human labellers every article a person has labelled | not yet measured | no labelled articles yet | no date recorded |
What this would look like if it were brokenA model that had learned nothing about framing would land on the same label as a person about as often as guessing does — around one time in five across five labels. A figure near 100% on a small sample is a warning too: it usually means the labellers saw the model's answer first. | |||
Survives an entity swap the most recent symmetry run | not yet measured | no symmetry test has been run | no date recorded |
What this would look like if it were brokenAn instrument that measured WHO an article is about rather than HOW it is written would keep its score when the sides were swapped, so this figure would collapse towards 0%. A perfect 100% on a handful of articles means the swap map matched almost nothing, not that the instrument is perfect. | |||
Stance quotes found verbatim in the text every analysis that returned a stance | not yet measured | no quotes have been checked yet | no date recorded |
What this would look like if it were brokenA model inventing its evidence would push this towards 0%, and every one of those stances is discarded rather than published — so the number you are reading is the share that survived, not the share we kept. It cannot be gamed upwards by relaxing the check, because the check is character-for-character against the exact text the model was shown. | |||
Agreement with a second model the most recent second-model run | not yet measured | no second-model run yet | no date recorded |
What this would look like if it were brokenTwo models that disagreed constantly would drive this towards chance, which would tell you the task is under-specified rather than that either model is wrong. We never average the two: a blended score is a number neither model would defend. | |||
Archive scored by the current method analysis method v7 | not yet measured | 0 of 0 stored analyses | no date recorded |
What this would look like if it were brokenThis is the denominator for everything above it. A quality figure measured over 4% of the archive is a different claim from the same figure over 90% of it, and the older articles are being re-scored rather than quietly excluded. | |||
Every figure is read from the same database that produces the scores, on the same page load. None of it is written down here by hand, and none of it is rounded in our favour. The sections below explain what each measurement is and how it is taken.
What it cannot tell you: these are aggregates, and every one of them is taken over a sample small enough that the sample size is printed beside it. A good score here does not mean the article you are reading was scored correctly.
Where our numbers come from
Everything on this site is built from four kinds of input. They are not interchangeable, and any figure we publish says which of them it was built from.
News coverage. Articles we collect from a fixed list of outlets and analyse one at a time. This supports claims about what the press raised, how it framed it, and who it took positions on. It supports no claim about anyone outside a newsroom. Its limits: we read the outlets on our list and no others, so a story only a local paper carried is invisible to us; and coverage volume reflects editorial choices, so an issue that is covered heavily is an issue editors chose, not necessarily an issue that grew.
How much of those outlets we capture. Not yet measured: the comparison against GDELT’s index has not been run for a published window.
Reader submissions. Text people send us on purpose, through a form that says what it is for, and which a person reads and approves before it counts towards anything. This supports claims about what readers chose to write in about, in their own words. It is not a sample of anybody. There is no sampling frame, and we make no attempt to build one. People who write in are self-selected, and the further limit is that a submission only counts once a human has approved it — so what is measured is a moderated stream, not a raw one.
Official records. Published government aggregates — grievance counts, administrative figures — which we use only to check our own numbers against something we did not produce. We do not currently ingest any such series, so this stream feeds nothing on this site today. Where a comparison has not been made, the category is marked unmeasured rather than being quietly grouped with the ones that failed a check — those are different facts.
Instrument records. What a measuring instrument recorded in a district and a week — today, satellite fire detections from NASA FIRMS; rainfall and reservoir storage only once a source for them has passed our own verification. This supports claims about what an instrument recorded, where and when. It says nothing about whether anyone needed anything, or whether the press should have covered it. It is switched off by default, shown only to our own analysts, and always drawn in its own frame beside coverage — never folded into a coverage figure, and never turned into a score. A detection that falls inside no constituency boundary is counted for the region and never assigned to the nearest place.
We never blend the streams into one figure without saying so. A count built from coverage alone can be one press conference reported eleven times; the same count with reader submissions beside it is a different and stronger thing, and you should be able to see which you have without asking us.
What we measure
We measure language and framing — which words an article chooses, whose account it leads with, what it treats as given, and what it leaves out. We call the left / center / right split The Spectrum: an AI estimate of how those framing cues distribute across an article, on a scale from 0 to 100.
That is a claim about the writing, not about the world. An article can be scored “center” and still be wrong, and an article scored “left” or “right” can be entirely accurate.
What we do not claim
- We do not fact-check. Nothing here says an article is true.
- We do not rate outlets. Every score is computed from one article’s text, and we never fold an outlet’s reputation into it.
- We do not claim objectivity. This is an AI estimate under a published rubric, and it can be wrong on any given article.
- We do not tune scores for anyone. There is no per-entity adjustment, no favourability setting, and no way to edit a published number by hand.
The active rubric
The rubric is the exact instruction the model is given. It is named, versioned, stamped on every analysis we store, and quoted in full below. If it ever changes, that change is visible here before it is visible anywhere else.
Currently active: Standard v1 (standard-v1)
You are a neutral political framing analyst. Your job is to analyze a news article and return structured analysis JSON. Do not add any other text. Rules: - Use ONLY the article text as evidence. Never infer framing from the news source name. - summary: 2-3 neutral sentences summarizing the article's key facts. - category: the single best-fitting top-level category. Choose EXACTLY ONE of: "Sports", "Politics", "Technology", "Business", "World", "Entertainment", "Health", "Science", "Other". Use "World" for international/general news, "Other" only when nothing else fits. - topics: 1-5 short, specific free-form topic tags for browsing/search (e.g. "IPL", "World Cup", "Social Media", "Arsenal FC", "Artificial Intelligence"). Use proper display casing. Return an empty array only if no specific topic applies. - sentimentScore: float from -1 (very negative) to 1 (very positive). - sentimentLabel: "positive", "neutral", or "negative". - leftPercentage, centerPercentage, rightPercentage: estimate the % of framing cues that lean left, center, or right. Must sum to exactly 100. - politicalFramingLabel: the dominant label. Use "left" if leftPercentage is highest, "right" if rightPercentage is highest, "center" if centerPercentage is highest, "mixed" if left and right are within 10 points of each other and both > 20, "unclear" if the article is factual/objective with little political language. - confidence: your confidence in the framing estimate (0-1). Use low confidence for objective news or when evidence is weak. - framingNotes: 2-4 bullet strings explaining specific language, word choices, or rhetorical patterns you observed. - loadedTerms: emotionally or politically charged words or phrases found in the text. Empty array if none. - disclaimer: must include "This is an AI-estimated analysis and may not reflect the article's actual intent or the publisher's editorial position."
Source blinding
A Blind Read is a score produced without the model ever seeing who published the article — the badge on an article page means that article was scored this way.
Before an article reaches the model we remove the publisher’s identity from it: the outlet name and its aliases, the domain, the byline, the masthead and the subscription boilerplate. Each removed span is replaced with the placeholder [PUBLISHER], and the model is told explicitly that the placeholder carries no information.
The point is structural. An instruction not to consider the publisher is unenforceable while the publisher is written across the text; removing it means the score cannot be a prior about who published the piece.
No blinded excerpt is available to show yet.
The symmetry test
If our score measured framing, then swapping the political sides in an article — every mention of one party for its rival, in English and in Telugu — should mirror the score: +0.3 should become −0.3. If instead the score barely moves, we are measuring who is being written about rather than how.
So we do exactly that, on a sample of real articles, and publish how far off the mirror we land. We call it symmetry error: score(swapped) + score(original), where 0 is perfect.
We have not run a symmetry test yet, so there is no number to show you here. When we have one, it will appear on this page whether it flatters us or not.
Stance, and the quote behind it
Alongside the framing split we record, for each actor an article takes a position on, whether the article’s language treats them favourably, unfavourably or neutrally — and the sentence that reading came from.
Every stance carries a quote, and the quote is checked character-for-character against the exact text our model was shown before the stance is stored. If it cannot be found there, the stance is discarded rather than published. That is what makes a stance something you can check instead of something we assert.
We run the same swap test on stance. Turning one party into its rival should carry the stance across unchanged — “critical of A” becomes “critical of B”. A stance that flips when only the name changed is the clearest evidence of the bias this whole exercise looks for, so we publish how often it happens.
Party alignment, and why we retired left and right
Left, centre and right is an American axis. In Andhra Pradesh and Telangana it measures very little: the contest is between named parties whose economic positions are close and whose regional, welfare and caste framings are not, and a reader asking whether an outlet is soft on the governing party learns nothing from a centre score of 62.
So we now measure party alignment: how much of an article’s coverage tilts toward each named party, as shares that add up to 100. “No party tilt” is one of the answers, and it is the most common one — most reporting tilts toward nobody, and a system that cannot say so would invent a lean for every routine story.
We kept the old left/right numbers in the database rather than deleting them, so figures published before this change stay comparable with figures published after it. They are no longer shown to readers.
This measures coverage, not parties and not voters. That an article tilts toward a party is a statement about the article’s language. It is not a claim that the party is right, that the outlet endorses it, or that anyone reading it agrees. We publish no candidate score, no seat projection and no electability figure, and we will not.
Beside the estimate we show who owns the outlet, where that is publicly documented, with a link to the document it came from. That fact is not produced by a model: a person checks the evidence and signs the row, and rows we could not verify are shown as unverified rather than guessed. The ownership and the tilt are two separate facts placed side by side. We draw no conclusion from the pair, because the data supports the two facts and does not support a third.
We also record whether an article frames its subject in caste terms, with the sentence that reading came from. No caste is ever named, inferred, stored or attributed to anyone. There is no column in our database that could hold one. This measures a property of the coverage’s language — the same way the policy frames below do — and if the sentence behind it cannot be located in the text we were given, the reading is discarded rather than published.
The swap test applies here too, and it is the sharpest version of it. Rename every mention of one party to its rival and the alignment should follow — an article that tilted toward A must come back tilting toward B. One that still reads A is being scored from what our model expects of coverage about A, not from the sentences in front of it. The caste-framing reading is tested the other way: it should not move at all, because renaming a politician does not change whether an article’s language is caste-framed.
How the story is framed
We also record which policy frames an article uses — whether it argues about cost, about legality, about public safety, about party advantage, and so on — as weights that add up to 100.
The categories are not ours. They are the fifteen frames of the Policy Frames Codebook (Boydstun, Card, Gross, Resnik and Smith, 2014), a published taxonomy used in political communication research. We use it unchanged so that the frames on this site mean what they mean in the literature, and so a frame count from last year is comparable with one from today.
Where a story is about
We also record which places an article is about — a state, a district, an assembly constituency — so that coverage can be counted by place.
This measures media coverage about a place. It does not measure public opinion held by people in that place. A district whose coverage leans one way tells you how the press wrote about it, and nothing at all about how anyone living there votes, thinks or feels. The two get conflated easily and quickly, so every place figure we publish is labelled coverage sentiment, never opinion.
The model names the places an article mentions, in the article’s own words. Turning a name into a place is done separately, by a fixed rule rather than by the model: an exact match wins, then a name that only one place in our list answers to, then a name plus a district or state the same article mentions. When a name could be two places and the article gives us nothing to tell them apart, we record it as ambiguous and it counts towards nothing. Indian place names collide heavily across states, and several constituency names are also common surnames — a wrong district is worse than a missing one.
We separate the place a story is about from places it merely mentions, and only the former counts towards a place’s coverage. A dateline is not a subject.
Our place list covers Andhra Pradesh and Telangana — two states, their districts, and all 294 assembly constituencies — with English and Telugu names for each. It is built from the Election Commission’s constituency numbering as published on Wikipedia and from Wikidata, both openly licensed.
Below the district, and Robinson’s problem
The gazetteer now goes below the district, to the mandal and to the units a corporator or a sarpanch is elected from. Smaller units make one particular misreading much more tempting, and it is a misreading with a name and a date.
In 1950 W.S. Robinson showed that a correlation computed over areas is not an estimate of the same correlation over people. Using 1930 US Census data he found percent Black and percent illiterate correlated at 0.77 across the 48 states and 0.95 across the nine census divisions — the same people, the same two variables, and a coefficient that moved with whichever aggregation level he happened to pick.
So “coverage of this ward is hostile to someone, therefore people in this ward are” is not a strong inference from our data. It is not an inference from our data at all. No sample size fixes this and no better model fixes it. Every surface below the district carries that sentence above its first number and writes it as the first line of any export, and every figure is labelled with the unit it was computed at — because per Robinson, a number without its aggregation level cannot be interpreted even when nobody is inferring anything about individuals.
Why we do not publish a “top 20 hotspots”
Test 294 assembly constituencies for “is coverage unusual here?” at the conventional 5% threshold, and even if nothing is unusual anywhere, about 15 of them come back significant by chance alone. For the 61 districts it is about three. Run it monthly and you get a fresh set of fifteen imaginary hotspots every month.
Any surface here that ranks or flags areas against each other applies Benjamini–Hochberg false-discovery-rate control and prints the Q it controlled at. A surface that cannot do that does not rank — it lists, with counts. What we do not yet do is account for neighbouring areas not being independent tests; the spatially-aware variants of the correction exist and we do not currently apply them, which is a limitation of our figures rather than a detail.
An area with no coverage is a result
Most mandals will have very few articles, and many will have none. We report “no local coverage found” distinctly from “outside the outlets we read”, and both distinctly from a blank. The first is a measurement: fifteen years of reporting across 200+ local papers (Hayes & Lawless, News Hole, 2021) find that local news volume, not slant, is what tracks civic engagement — so an area our outlets never write about is a finding about that area, not a hole in our data.
Coverage of a constituency is not a poll
We produce a per-constituency brief for clients: how much coverage a seat is getting, which actors that coverage is for or against, which policy frames drive it, which outlets differ, and what has moved since the previous period.
Every figure in it describes press coverage. None of it measures public opinion, support for a candidate, or how a seat is likely to vote. An actor covered favourably in a seat is an actor the press wrote about favourably there — that is the whole of the claim. We publish no single summary score for a seat or a candidate, deliberately and in no version of the product: one number is precisely what invites being read as a prediction.
Positions on an actor come from the same single analysis pass as everything else, and each one carries the verbatim sentence it was scored on. A position whose quote we cannot locate in the text the model actually read is discarded rather than stored. Positions are only counted for actors we can resolve to a known name — a raw string would merge two spellings of one person and split one person into two.
Policy frames use Boydstun et al.’s Policy Frames Codebook (2014), a published taxonomy, rather than categories a model invented for us.
Where a constituency, an outlet or an actor has too little coverage in the period, we compute no average for it at all: the cell shows the number of articles and says so. “Not enough coverage to say” is a different statement from “balanced”, and we would rather show an empty cell than a flattering figure built on three articles.
Who pays for a brief, and what a paying client can and cannot do with it, is answered on the who we are page (PRD-003) — ownership, funding, our client categories, and the conflicts register.
What people are asking for is not a survey
We record, from each article and from reader submissions people send us on purpose, what is being asked for: which issue, what kind of problem (a shortage, a price, a delay, a safety risk), which place, and the sentence that says so. We also record the other half — a sanctioned scheme, a tender, a completed work, an official reply — so that a demand can be shown beside whatever answer it did or did not get.
This measures need that is expressed in news coverage and reader submissions. It is not a survey and does not measure the views of everyone in a place. There is no sampling frame here and no attempt at one. People who write to a newspaper, and journalists who choose what to cover, are not a random sample of a district. A cell that is loud is a cell that was written and submitted about often — that is the whole of the claim, and it is a different claim from what a district wants.
We report four things separately and never collapse them into one score: how often something is raised relative to how much that place is covered at all, how severe the signals are, how many consecutive weeks it has been raised, and whether anything has been recorded in reply. A single index would be easier to read and impossible to check — a reader who cannot see which part moved cannot argue with it.
For “is this new” we use a published method rather than a threshold of our own: Kleinberg’s burst detection (2003), which distinguishes a one-week spike from a sustained change in rate. A single loud week is a story; six steady weeks is something else, and treating the two the same is how a dashboard becomes noise.
Where a place and an issue have too few signals in the period, we compute nothing at all for it: the cell shows the count and says it is below the floor. “Not enough signals to say” is a different statement from “nothing is needed here”. The floor for a needs figure is higher than the one for a coverage count, because a claim about what a place needs is a stronger claim than a count of articles.
Every figure also shows what it is made of — coverage, reader submissions, or both. A number built only from news coverage means something different from one that people independently wrote in about, and you should be able to see which you are looking at without asking.
These signals come from the same model call that produces the summary and the framing split, rather than from a separate pass. If that ever changes, this paragraph changes with it: which instrument produced a number is part of the number. Whatever the extractor, the checks that decide whether a signal is kept — the confidence floor, and the requirement that its quote be locatable word-for-word in the text the model was shown — run in our own code and are never left to the model.
We have not yet compared these figures against an independent series. Until we have, no category on this site is described as validated — and this paragraph is what we show instead of a number.
We run the same swap test on needs. Nothing about a water shortage depends on which politician is quoted beside it, so a need that appears or vanishes when only the names change is the extractor reading party cues instead of stated problems.
Three measurements that only exist across the corpus
Narrative lead-lag
This measures predictive precedence between two outlets' coverage of one frame: whether one outlet's coverage helps predict the other's. It is not a measure of influence, and two outlets covering the same event will show precedence without either affecting the other.
For one event and one policy frame, we build a daily series of framing intensity per outlet — including the days with none, because a series that omits its gaps turns three scattered mentions into three consecutive ones. We test each series for a unit root (augmented Dickey–Fuller), choose a lag by AIC over a maximum of seven days, fit a vector autoregression, and run the Granger Wald test. We then check the residuals for leftover structure (Ljung–Box) and for normality (Jarque–Bera), and we publish those diagnostics beside the result: a model whose residuals fail is reported as failed rather than quietly published.
Every outlet pair crossed with every frame is one family of tests. Left uncorrected, at the conventional 5% threshold, a family that size returns a significant result every single week by chance alone. We correct with Benjamini–Hochberg across the whole family, and the corrected value is the only one any surface shows. Below 60 days of coverage from both outlets we do not test at all, and the cell says so with its count.
The limitation: Granger precedence is not causation and this method cannot distinguish the two. Two newsrooms at the same press conference will show precedence with neither affecting the other.
Coverage deserts, and the news desert atlas
This measures news coverage of a place in the outlets we read, per head of population. A low figure means little was written about that place; it does not mean little happened there, and it is not a judgement about the place or the people who live in it.
We work at mandal level, and we report Napoli’s three counts separately — every story that mentions a mandal, the subset datelined there, and the subset in a critical-information category. They are three numbers because they mean three things: forty syndicated national stories mentioning a place is not local news, and any judgement about local coverage uses the datelined count. There is no fourth number combining them, in the console, on the public atlas, or as an export column.
A zero is only publishable if we looked. Every district listing page we fetch records when it last returned a story — on content, never on an HTTP status, because two of the sites we read answer an unknown address with their own front page. A mandal whose district listing was not reached during the window renders as not measured: a third state, distinct from both covered and low-coverage. Reporting “we did not look” as “there is no news there” is the one error that would make this figure worthless, so it has its own state, its own texture and its own sentence, and it is named first in every legend rather than last.
We publish no coverage-per-capita figure, and this is the reason. No mandal is classified as a desert in this edition. That judgement is a threshold over stories per head of population, and per-constituency population is not published in India — the census and delimitation figures are per district, and deriving one from the other is a modelling assumption, not a measurement. This atlas therefore publishes counts and coverage states, never coverage per capita. The low-coverage threshold in the definition is a floor over the distribution of the mandals we measured — not a model and not a score — and because it needs that denominator, no mandal currently carries it. The class is named in the legend anyway: a legend with two entries would imply the third state does not exist, when what is true is that we cannot yet reach the judgement.
Nothing is ranked. Nothing here is ranked. Ordering 1,309 mandals by a count invites a top-20 reading, and about 65 of them would test unusual by chance alone at the conventional threshold. We have no per-mandal significance test to correct with, so this is a list with its counts and its order is alphabetical. Where we do rank across units elsewhere, we control the false discovery rate with Benjamini–Hochberg and print the corrected count; that implementation is the standard step-up and is not spatially aware, which matters because neighbouring mandals are not independent tests.
The atlas is public at /atlas, with the whole dataset downloadable as CSV and GeoJSON. Creative Commons Attribution 4.0 (CC BY 4.0). Attribute to Vekuva, naming the window and the snapshot this file carries. The GeoJSON carries no geometry, deliberately: we hold no mandal boundaries, and a mandal is a sibling of an assembly constituency rather than a part of one, so there is no polygon we could honestly attach. Each feature carries its LGD code instead, which is the join key every Indian government dataset already uses.
The limitations: we measure the outlets we read, which is not every outlet; a mandal with no population figure gets no rate rather than an assumed one; and this is a measure of coverage, never a judgement about a place or the people who live there.
Coverage–vote divergence
This measures how press coverage in the 90 days before a poll differed from the recorded result in the same place. It does not measure persuasion, influence, or what any voter thought, and it is not a forecast: the election it compares to has already happened.
For one mandal we compare two vectors: the party mix of coverage in the 90 days before a poll, and the recorded vote share in the polling stations we could link to that mandal. We publish both vectors and their difference in percentage points, per party, with a 95% Wilson interval on the coverage share. There is no single divergence number — no swing index, no rating — in the console, in an export column or behind a flag.
Booth results come from the state Election Commission’s own Form 20 publications. Every constituency we publish has passed two arithmetic checks: each station’s candidate columns sum to the total printed on the form, and the station totals sum to the published constituency total. A constituency that fails either is reported unparsed and named, never force-committed. Stations are matched to villages by deterministic string matching with a recorded score and a human confirmation step; stations we could not place are counted, published as unplaced, and never dropped, because the hardest stations to place are disproportionately rural. Their share is printed beside every figure.
The limitations: the smallest unit we ever display is the mandal, even though the data goes finer; the comparison is retrospective and describes a contest that has already been decided; and a difference between coverage and a result is a difference, not a mechanism.
The derived measures, and what would show they are wrong
Beyond the per-article analysis we compute a catalogue of derived measures over the whole archive — how concentrated a place’s coverage is, whether a story split the press into two camps, which outlet reported first, and whether Telugu and English coverage of the same event framed it differently.
Every one of them is published here with four things: how it is computed, the minimum sample below which we compute nothing at all, what it cannot tell you, and the scenario that would show the implementation is wrong. That last one is unusual to publish and it is the point: a measure nobody can falsify is decoration, and each of these scenarios is a test in our own test suite.
None of these is marked validated. That word is reserved for a measure correlated against an independent series that we did not produce, and none of them has been. They are published as unvalidated or experimental instead. Two more are published with no figure at all, because the data they need does not exist — per-constituency population is not published in India, and we have not ingested an official grievance series. We would rather show you a defined measure we cannot yet compute than an approximation of it.
Coverage concentration
index (0–1) · tier 1Sum of the squared shares of a place or topic's articles held by each source in the window (the Herfindahl–Hirschman index). One outlet publishing everything scores 1; ten outlets publishing equally score 0.1.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 10 articles. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- A place covered equally by ten outlets returns above 0.2 — that is 1/10 plus rounding, so anything higher means the shares are not shares.
- What it cannot tell you
- It measures who published, not who was read. A dominant outlet with no audience scores identically to a dominant outlet everybody reads.
Framing polarisation
separation (0–1) with standard deviation · tier 1Standard deviation of members' bias_score, plus a separation figure: the largest gap between consecutive sorted scores as a share of the total range. Above half the range, with at least two articles either side, the cluster is reported as two camps rather than one spread.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 5 articles in the cluster and 3 distinct sources. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- A cluster where every outlet scores +0.5 reports high polarisation — identical scores have zero spread, and a bimodality figure computed from them is reading noise.
- What it cannot tell you
- Two camps in framing is not two camps in fact — an event can be genuinely two-sided and score exactly as one the outlets simply disagree about. The separation figure also detects exactly two camps: a three-way split is reported by its widest gap. Sarle's bimodality coefficient was rejected for this metric because its sample-size correction makes it unable to exceed its own threshold below roughly fifteen members, and clusters here hold five to twenty.
Blindspot rate
rate (0–1) with denominator · tier 1For each outlet: the share of major events — those covered by at least the qualifying number of distinct sources — that the outlet published nothing about, within the window.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 20 qualifying events. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- An outlet that publishes on every qualifying event shows a rate above zero, or an outlet absent from the archive entirely shows a rate at all rather than being excluded.
- What it cannot tell you
- Not covering a story is not the same as suppressing it. A specialist outlet legitimately skips most general news, and this metric cannot tell the two apart — which is why the badge wording states a count and never a motive (§36).
Story lifespan
hours · tier 1Hours from an event's first_seen_at until half its eventual articles had published (the half-life), plus the total span from first to last article.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 6 articles in the cluster. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- A story whose articles all published within one hour reports a long half-life — the half-life cannot exceed the total span.
- What it cannot tell you
- It measures the archive's coverage of a story, not the story. An event this product's sources ignored for two days has a lifespan starting when they noticed.
First-report credit
events led, with median hours · tier 1Per outlet: the number of events where it published the earliest article, and its median lead time in hours over the second outlet to publish.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 10 events. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- An outlet that published after another on every event is ever credited with a lead, or a lead time comes back negative.
- What it cannot tell you
- Leading in this archive is not leading in the market. An outlet this product does not scrape cannot be beaten, and a story broken on television arrives here when somebody writes it up.
Agenda lead–lag
correlation (−1 to 1) at a lag in weeks · tier 1Pearson cross-correlation between two outlets' weekly article counts on a topic, at lags of 0 to 3 weeks. The lag with the highest correlation says who moves first; the sign of the lag says which way.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 12 weeks of series and 5 articles per week on average. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- An outlet is reported as both leading and following the same partner on the same topic at the same lag — a relationship that cannot hold in both directions at once.
- What it cannot tell you
- Correlation at a lag is not causation, and both outlets may be following a third party neither of them is. It says the series move together in an order, nothing more.
Numeric discrepancy
ratio (≥1), flagged above tolerance · tier 2Per event and quantity: the ratio of the largest to the smallest value reported for the same quantity in the same unit. Flagged when that ratio exceeds the tolerance and the values are not a rounding of each other.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 3 outlets reporting the same quantity. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- It flags 1.2 crore against 1.24 crore as a discrepancy, or fails to flag 40 deaths against 400.
- What it cannot tell you
- Outlets can legitimately report different figures for the same quantity at different times as a toll rises. A flag is a prompt to look, never a finding of error.
Stance volatility
standard deviation, with dated flips · tier 2Per outlet and entity: the standard deviation of signed stance over the window, plus each date where the sign of the trailing mean changed and stayed changed for the confirmation window.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 8 stances for the outlet–entity pair. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- A single low-confidence stance against the run of the series triggers a flip.
- What it cannot tell you
- An outlet's stance changing is not an outlet changing its mind — the story may have changed. The metric dates the change; it does not explain it.
Claim propagation
outlets, with span and contradiction count · tier 2Per claim: the number of distinct outlets that published it, the span in hours from first to last, and the number that published a contradicting claim, all at the existing adjudication confidence floor.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 2 outlets carrying the claim. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- Propagation exceeds the number of articles containing the claim — a claim cannot reach more outlets than it appears in.
- What it cannot tell you
- Repetition is not endorsement: an outlet quoting a claim to dispute it is counted as carrying it unless the contradiction was itself adjudicated.
Source diversity
distinct attributed sources · tier 2Per article: the count of distinct attributed sources quoted, split by kind — named person, named official, document, and unnamed attribution such as 'sources said'.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 1 article. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- An opinion column with no attribution at all scores above zero, or the same person quoted four times counts as four sources.
- What it cannot tell you
- Counting attributions measures reporting effort, not accuracy. A piece citing six sources who all say the same wrong thing scores well.
Coverage–demand gap
share difference (−1 to 1) · tier 2Per place: reader attention share (views plus searches within the window) minus coverage share (articles within the window). Both are shares of their own within-window total, so each sums to 1 across places.
- Status
- Experimental — the definition may still change
- Minimum sample
- Needs at least 20 views and 10 articles. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- Either share fails to sum to 1 across the window — that means a denominator was computed per row rather than across the set.
- What it cannot tell you
- Readers cannot demand coverage of a story nobody published. Attention is measured on what exists, so a genuine blindspot can register as no demand at all.
Coverage per capita
share difference (−1 to 1) · tier 3Per place: share of the window's articles minus share of the population. Negative means a place is covered less than its population share — a coverage desert.
- Status
- Not computed — the input data does not exist yet
- Minimum sample
- Needs at least 10 articles. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- It reports a per-constituency figure at all while the only available population data is per district.
- What it cannot tell you
- Population share is not news-worthiness share. A district with a capital city legitimately generates more news per head, and this metric will always mark rural districts as under-covered.
- Why there is no figure
- Per-constituency population. Census and delimitation figures are published per district, not per assembly constituency; deriving one from the other is a modelling assumption, not a measurement.
Grievance–coverage correspondence
Spearman rho (−1 to 1) · tier 3Per district and issue: Spearman correlation between official grievance volumes and coverage volume, week by week. Divergence is either a media blindspot or a grievance-reporting gap.
- Status
- Not computed — the input data does not exist yet
- Minimum sample
- Needs at least 12 weeks with both series present. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- It reports a correlation computed against this product's own reader submissions rather than an independent official series — that is the archive correlated with itself.
- What it cannot tell you
- Neither series is a count of problems. Grievances measure who filed, coverage measures who wrote; they can diverge for reasons that have nothing to do with need.
- Why there is no figure
- Official grievance volumes. The ingestion described in prompts/civic-intelligence/12-verification-spikes.md (Spike 1) is not built, so there is no independent series to correlate against.
Language framing divergence
difference in framing score (−2 to 2) · tier 1For events covered in both Telugu and English: the difference between the two language sets' mean framing split, mean sentiment, and whether their dominant policy frame agrees.
- Status
- Unvalidated — computed, not yet checked against anything independent
- Minimum sample
- Needs at least 3 articles in each language, for the same event. Below it we compute nothing — there is no figure to withhold, because none was produced.
- You would know it is broken if
- Divergence correlates with how many articles an event has rather than with language — a difference that grows with sample size is measuring sampling noise, not a language effect.
- What it cannot tell you
- Language is detected from the text, not declared by the outlet (§31). An English wire story on a Telugu masthead counts as English, which is correct here and surprising if you expect outlet-level buckets.
Agreement with human labellers
Separately, people read a sample of articles and give their own left / center / right split without seeing the model’s answer first. Comparing the two tells us whether the score tracks what a careful reader would say.
No articles have been hand-labelled yet, so there is nothing to report here. We would rather show you this sentence than a number we have not earned.
A number on its own is not a claim anyone can check, so the table below names the thresholds we are measuring against and says which band our figure falls in. The bands are the conventional ones from the inter-rater reliability literature, not thresholds we chose after seeing our own result.
| Band | Coefficient | What it permits |
|---|---|---|
| Reliable | ≥ 0.80 | Conclusions may be drawn from this field directly. |
| Tentative | 0.667 – 0.80 | Tentative conclusions only. Usable as a signal, not as a headline number. |
| Insufficient | < 0.667 | Not reliable enough to draw conclusions from. Published anyway, because hiding it would misrepresent the fields that do clear the bar. |
No hand-labelled articles yet, so there is no coefficient to place in a band. We would rather show you this sentence than a number we have not earned. Bands are Krippendorff's conventional cut-offs: 0.80 and above is reliable, 0.667 to 0.80 supports tentative conclusions only, and below 0.667 is insufficient to draw conclusions from. The coefficient we publish is Cohen's κ, computed over two raters on a nominal label. On that design κ and Krippendorff's α coincide; they diverge once there is a third labeller or an unanswered field, and this page will say so when that happens.
Our frames, and where they come from
The fifteen policy frames we report are not a list a model proposed. They are the Policy Frames Codebook of Boydstun, Card, Gross, Resnik and Smith — a published taxonomy designed to be near-exhaustive across any policy issue, and used across political communication research. We adopted it whole rather than inventing categories and mapping them onto it afterwards.
That means this table is a correspondence, not a translation: the two columns are the same fifteen dimensions. It is here because four of our stored labels are shortened, and a researcher comparing our output to the Media Frames Corpus needs to know that our “Public sentiment” is the codebook’s “Public Opinion” and not a category of our own.
| Our label | Policy Frames Codebook | Stored key |
|---|---|---|
| Economic | Economic | economic |
| Capacity and resources | Capacity and Resources | capacity_and_resources |
| Morality | Morality | morality |
| Fairness and equality | Fairness and Equality | fairness_and_equality |
| Legality and constitutionality | Legality, Constitutionality and Jurisprudence(we shorten this) | legality_and_constitutionality |
| Policy prescription | Policy Prescription and Evaluation(we shorten this) | policy_prescription |
| Crime and punishment | Crime and Punishment | crime_and_punishment |
| Security and defense | Security and Defense | security_and_defense |
| Health and safety | Health and Safety | health_and_safety |
| Quality of life | Quality of Life | quality_of_life |
| Cultural identity | Cultural Identity | cultural_identity |
| Public sentiment | Public Opinion(we shorten this) | public_sentiment |
| Political | Political | political |
| External regulation | External Regulation and Reputation(we shorten this) | external_regulation |
| Other | Other | other |
The list is frozen. Adding a frame would change the meaning of every distribution already stored, so it is a migration and a re-score rather than an edit — the same rule that governs every other change to how an article is scored.
A second model, and what we do not do with it
We periodically re-score a sample of articles with a different model, under the identical instruction and the identical output contract, and record where the two disagree.
We never average the two, and we never publish a blended score. Averaging two models produces a number neither model would defend, and it would mean the figure on an article page was no longer the output of one stated method. Disagreement is used to flag an article for review and to tell us which outlets and which topics the instrument is least reliable on — which is a different and more useful thing than a smoother number.
Two models agreeing does not make either one right. They can be wrong together, particularly on the kind of article that is genuinely contested. It is the human comparison above, not this one, that tests whether the score tracks what a careful reader would say.
Watching for drift
A model’s behaviour can change without anyone touching our code, and an outlet’s can change without anyone touching theirs. So we record what the instrument says about each source every day and look for a step change in it.
When we find one, a person looks at it. We do not adjust a published score to compensate: there is no code path in this product that edits a stored number, and a silent correction would make every figure on this page describe a system that no longer exists.
School exam results, and what a pass rate is not
We hold school-level board exam results for Andhra Pradesh and Telangana, joined to the government’s own school register (UDISE+) and rolled up to districts and constituencies. Every figure is a school aggregate.
There is no student record anywhere in this product, and no column that could hold one. Not a name, a hall ticket, a roll number, a date of birth or a mark. We do not ask for a hall ticket number and we do not pass one on: individual results are published by the examinations boards, and we link to them. That is enforced by the absence of the fields rather than by this paragraph.
The join
A school is placed by its mandal, matched from the register’s own block and village fields by a fixed rule — never by a model — with the match score recorded. Only a single exact match is written automatically; everything else is confirmed by a person. A school whose mandal we cannot resolve is published as unlinked and counted, never guessed into a neighbouring mandal and never quietly dropped: the schools hardest to place are disproportionately rural, so dropping them would bias every figure in one direction.
A mandal is not inside an assembly constituency — the two are drawn over the same ground by different departments. The only crosswalk between them we have is booth-level election data, which records both. Where that is not yet loaded for a seat, the page says so rather than showing zero: we cannot resolve this seat and this seat has no schools are different sentences.
The three things a cell can say
- A result was published — the board published a figure for this school.
- No candidate passed — the board published a figure and it was zero. That is a finding, and we publish it as one, by name.
- No result was published — the board published nothing for this school. That is an admission about the source, not a statement about the school. It is never counted as a failure and never enters a pass rate’s denominator.
The last two are separate values in the data, not one grey cell. Collapsing them would file a school nobody published a result for as a school that failed every child it entered.
The floors
A school with four candidates has a pass rate of 0, 25, 50, 75 or 100 per cent and nothing in between. Below ten candidates we publish the count and no rate — and the rate is withheld in the database, not merely left undrawn, so it is absent from the page’s data as well as from the screen. A school below the floor is also never named in the zero-per-cent list: a school of three that passed nobody is a small number, not a finding.
Lists of districts and constituencies are ordered alphabetically, never by pass rate. Comparing 294 constituencies at once produces about fifteen “significant” results by chance alone, so a ranking needs a multiple-comparisons correction — and a surface that does not apply one lists with counts instead of ranking.
The three things this cannot say
- “This school is better than that school.” A pass rate is not a school quality measure. It is an outcome measure confounded by intake — a school that enters every child it teaches is not comparable to one that enters the children it expects to pass. A league table of government schools also names teachers. There is no rank, no score and no comparison to a 'similar' school anywhere in this product, because a peer definition is something we would have to measure and have not (§35.3).
- “This MLA improved education.” The constituency roll-up is a fact about schools in a geography, not about a person's performance. Where the result sits beside how a mandal voted, the areas-not-people constraint is at the top of the page with no dismiss control.
- “These children did well or badly.” Every figure here is a school aggregate. There is no student record, no hall ticket, no mark, and no column that could hold one — the guarantee is the absent column, not a policy (§35.4).
Where it comes from
School records come from UDISE+ on the national open data platform; results come from the state examinations boards’ own publications. Every row carries the URL it came from and the date the source asserts. Nothing is loaded by a scheduled job: a person reviews the parse against the board’s own district totals and signs it, and an unsigned or arithmetically inconsistent file is refused rather than loaded with a warning.
Data sources and attribution
The reference data behind places, maps, elections, schools and our own quality checks comes from these sources, under the licences they are published with. Where no licence is recorded we say so rather than guess.
The news itself belongs to the outlets that published it. We read each article in full to analyse it, but we show readers only its opening — at most two paragraphs — with our own summary and analysis, and a link to read the whole story on the outlet’s site. We do not republish their photographs, and we place no ads inside their text.
- datameet — India assembly constituency boundaries — constituency and district boundaries on maps, and the positions of the hex tiles derived from them. Licence: CC BY 2.5 IN.
- Wikipedia — constituency lists and election results by constituency — constituency numbers, names and districts, and assembly election results. Licence: CC BY-SA 4.0.
- Wikidata — Telugu place labels, legislators' positions held, and proposed outlet ownership. Licence: CC0.
- Local Government Directory (LGD), Ministry of Panchayati Raj — mandal codes, names and their parent districts. Licence: GIGW copyright policy (lgdirectory.gov.in) — reproduce free of charge, identify the source.
- Election Commission of India — statistical reports — cited as the upstream source of constituency election results. Licence: GoI (open).
- Chief Electoral Officer, Andhra Pradesh — Form 20 booth results — polling-station results. Licence: Government of Andhra Pradesh — public record.
- Chief Electoral Officer, Telangana — Form 20 booth results — polling-station results. Licence: Government of Telangana — public record.
- MyNeta (Association for Democratic Reforms) — candidate affidavits — candidates' declared criminal cases, assets and education. Licence: see source; mirrored verbatim with attribution.
- UDISE+, Ministry of Education (via data.gov.in) — the school register. Licence: Government Open Data Licence – India (NDSAP), data.gov.in.
- Board of Secondary Education, Andhra Pradesh — school-level examination results. Licence: Public record of a state examinations board (bse.ap.gov.in).
- NASA FIRMS — active fire detections — the instrument series (fire detections by constituency), where enabled. Licence: NASA Earth science data: open, no restriction on use; citation requested.
- The GDELT Project — DOC 2.0 API — checking how completely we capture coverage (counts and domains only). Licence: GDELT Project terms of use — unrestricted use with citation.
Known limitations
- Framing is genuinely contested. Two thoughtful readers disagree on the same article, which puts a ceiling on any agreement number — including ours.
- Our swap map covers Indian national and Telugu state politics. Articles about other political systems are skipped by the symmetry test rather than measured by it.
- Blinding is thorough but not perfect: an outlet can be implied by a house style or a stringer network we do not redact.
- We read the first 8,000 characters of an article, not all of it. That is roughly 1,300 words, which covers the great majority of what we collect — but an actor who is first mentioned after that point cannot receive a stance, a place first named after it cannot be tagged, and a frame introduced only in a long article’s closing section may be missed. This is a property of scoring each article in a single pass, and we would rather state it than have it discovered.
- Samples are small. Every number on this page is published with its sample size for exactly that reason.
- The model can be wrong on a single article in ways no aggregate metric will show you.
Last measured
No symmetry measurement has been run yet.
Questions about any of this? Back to the news.