Imagine trying to grade the economic strategies of forty different countries for how seriously they take resource security, using documents that range from five-year development plans to budget speeches to national security reviews, written in a dozen languages, by governments with wildly different priorities and vocabularies. Now imagine doing that fairly, consistently, and without your own biases quietly creeping into the scorecard.
That’s the puzzle Global Footprint Network and Greenings set out to solve in their new report “Prepared for the Predictable? Overshoot and the Failure of Countries to Ready Themselves.” The solution? Deploying AI as independent evaluators to score a diverse set of national strategies against a standardized rubric.
How do you grade a strategy document?
It’s easy to read a strategy document and get a vague, gut-level impression: “this one feels serious about resources,” or “this one doesn’t seem to care.” But such impressions may be biased, misleading, and lacking any semblance of a scientific method. So, the Global Footprint Network and Greenings researchers did something a good teacher does before grading a stack of essays: they wrote down, in advance, exactly what both great and mediocre answers would look like, and what failing ones would look like too, in observable, specific terms.
They landed on two big questions worth asking of any country’s economic strategy. First: does it recognize overshoot as a real threat, not just a footnote, but something baked into how it thinks about its own economic future? Second: does it do anything about it? Does recognition translate into binding policy, funded programs, and enforcement, or does it stay a nice sentence in the introduction with no teeth behind it and little connection to its priorities? The researchers split these two questions into five specific dimensions, each with a five-point scale, from “fully blind to the risk” all the way to “fully integrated into the country’s core planning.” Each dimension was written with enough concrete detail such that two independent readers would draw roughly the same conclusions.
Distilling the Data
Before handing the documents over for scoring, the researchers had to solve a practical problem: national strategy documents vary significantly in length, with some being over 300 pages long. Imagine holding all three hundred pages of information in your head while simultaneously memorizing and reasoning through all the key elements and rules of scoring. Even current LLMs struggle to do this without making up answers.
To handle this volume, a single AI model was deployed as a first-pass extractor. It is important to note that this initial stage involved absolutely no scoring or interpretation. Instead, the AI was tasked purely with ingestion and summarization. It processed the lengthy raw documents and distilled them down to their core components: generating an executive summary, identifying key policy pillars, and extracting exact wording and evidence regarding the country’s resource priorities.
By isolating the concrete evidence from hundreds of pages of bureaucratic filler, the team ensured the subsequent grading phase would be focused on concentrated policy facts rather than getting lost in the noise.
A panel of AI judges
Instead of having a small team of humans read dozens of dense national strategy documents and apply the rubric themselves, which invites fatigue, inconsistency, well known biases, and a project timeline measured in years, the researchers turned the rubric into a structured prompt. This prompt was then deployed across four distinct Large Language Models built on entirely different architectures: Google’s Gemini Flash 3.5, OpenAI’s GPT-oss-120b, Tencent’s Hy3, and NVIDIA’s Nemotron Super.. Each model read the same documents and scored them against the same five-dimension rubric, blind to what the other models found, independent of sequence. Think of it as convening four independent judges for a competition, except each judge reads with perfect patience, never letting personal opinions or biases get in the way.
The result provided a crucial validation: the models largely agreed. Because these AI models were built by different developers and process data in fundamentally different ways, the convergence of scores is an important structural validation. It suggests that the rubric is doing its job—capturing objective, observable signals in the text rather than simply reflecting the quirks of a single model’s architecture.
The findings
The findings themselves are worth a headline of their own. China’s latest five-year plan came closest to the ideal of full recognition paired with full action, though even that was far from a perfect score. Switzerland’s various strategy documents told wildly different stories, depending on which one you read, While Argentina consistently scored at the bottom across every report analyzed. Notably, the European Union, despite years of headline-grabbing Green Deal initiatives and rigorous sustainability reporting rules, scored surprisingly low, a reminder that talking a lot about sustainability and doing something about it are not the same thing.
Opening the black box, what were the AI judges thinking?
Here’s the part that turns this from an interesting research exercise into something genuinely useful: the researchers didn’t just publish a chart and move on. They made the entire body of AI-generated analysis, every score and every piece of reasoning behind it, publicly queryable through Gemini Notebook, an AI powered research tool.
That means you don’t have to accept this report’s word for why Argentina scored low, or why Switzerland’s reports are so scattered, or why the US and Canada land in similar places for such different reasons. You can simply ask, and the AI tool will walk you through the reasoning behind the score, document by document.
What is next
The report even offers a starter list of intriguing questions to try: Why does China score closest to full preparedness? What’s driving France’s second report to score so much higher than its first? Why do resource-rich countries like Argentina, Chile, Uruguay, and Brazil end up scoring so differently from each other, despite starting from a similar resource-rich position? These aren’t rhetorical questions dropped in for flavor, they’re genuinely answerable, and the tool that answers them is sitting right there for anyone to use.
There’s something quietly optimistic in all of this. Using AI to grade something as consequential as national resilience could easily have gone the way of every AI-skeptic’s worst nightmare: opaque, unaccountable, take-it-or-leave-it, and perhaps full of hallucinations. Instead, this project leaned the other way, toward transparent criteria, multiple independent models cross-checking each other, and an open door for anyone to interrogate the reasoning themselves. It is used for consistent evaluation, not for unaccountable interpretations. That’s AI being used less like an oracle and more like well-organized, patient, indefatigable research assistants, who do not ask you to trust them, but hand you the whole package, from the rubric, the receipts, the red pen to the results, and maybe even point you toward new horizons.
This blog was authored by Mathis Wackernagel, Kristin Kostka, and David Lin
Additional Resources:
♥ The full, downloadable report
♥ Technical Methodology which describes the process of the evaluation and documents the prompts used
♥ A summary for policy analysts
♥ A more popular blog introducing the report
♥ A blog on the AI-based research methodology behind the report [this one here]
♥ A blog on the report’s conclusions
♥ Access to a Gemini Notebook where you can query the LLM outputs for each country
♥ Access to a Gemini Notebook where you can query all the information embedded in more context, as the output is enriched with the researchers’ previous publications and perspectives
♥ Earth Overshoot Day 2026 press release
♥ Prepared for the Predictable? press release