โ† Back to playground

RESEARCH / AN EXPERIMENT / MONGOLIA

How far can you take research you already have?

I wrote three papers about Mongolia as a student: one on mining and school enrolment across the provinces, one on what can be done about an economy built on mining, and one on the social cost of the 2019 raw coal ban. Five years later I handed all three to AI, along with the datasets and the regressions, to find out what the work could become. This page is about what turned out to be possible, not about whether I had been right.

01 / Where the papers stopped

Three papers that never spoke to each other

Each one was written for a different class and answered a different question. The first compared Mongolia's twenty one provinces between 2000 and 2022 and found that large-scale industrial mining and informal small-scale gold mining pull school enrolment in opposite directions. The second asked what can be done about an economy that cannot simply stop mining. The third looked at Ulaanbaatar's 2019 raw coal ban and found that the air got cleaner while carbon monoxide poisonings rose among the poorest households.

They are obviously about the same country and the same people. A family in a ger district is in all three. But each one was built from one dataset, with one method, inside one deadline, and none of them could reach the other two. They stopped when the quarter ended.

02 / What became possible

The same material, and a different ceiling

The interesting question was never whether an AI would agree with me. It was what a person with three finished papers and no institution behind them can now actually build out of them.

Comparison table of what the three papers could do and what the same material can do now
WHAT THE PAPERS COULD DO, AND WHAT THE SAME MATERIAL CAN DO NOW

Four of those rows matter more than the rest. The data came out of the statistics office's database rather than off a screen and into a spreadsheet, so the panel runs to 2025 instead of stopping in 2022, and can be re-run whenever those tables update. More than one source is in the same room now: World Bank national accounts, EITI revenue figures, FAO livestock counts, ILO estimates, UN projections and PISA scores sit alongside the provincial tables, and where two of them disagree, the disagreement is shown rather than smoothed over.

The three papers became one argument, because mining, schooling, air pollution and child labour are happening to the same households. And the work can now ask a forward question. The papers described what had happened. The report asks what it implies for a country whose herd fell from 71.1 million animals in 2022 to 58.1 million in 2025, where a dzud that used to come once a decade now comes one year in five, and whose coal revenue depends on a Chinese steel industry in decline.

03 / Where AI actually fits

It is not evenly distributed

The more useful thing to come out of this is not the report. It is a map of where a tool like this changes research and where it does not, drawn from actually doing it rather than guessing.

Chart of seven research stages showing what AI carried and what stayed with me
SEVEN STAGES OF ONE RESEARCH PROJECT, AND WHO CARRIED EACH ONE

The middle is where it changes everything. Finding the series, pulling them, aligning them, and working out which of four sources to believe is most of the actual labour of research, and it is now close to free. That was also the part that made me stop the first time.

The two ends barely moved. Choosing what to care about is still the whole of the beginning, and deciding what I am willing to claim is still the whole of the end. AI argued against my conclusions, which was genuinely useful, but somebody still has to decide which of those arguments to accept.

04 / A side effect worth keeping

Everything got checked along the way

This was not the point of the exercise, but it is a real benefit of it, so it belongs on the page. Two things happened while the analysis was being extended. The findings got tested against data they had never seen, and every quantitative claim in all three papers got traced back to a primary source.

Table of the first paper's eight findings checked against provincial data for 2023 to 2025
THE FIRST PAPER'S EIGHT FINDINGS, CHECKED AGAINST PROVINCIAL DATA FOR 2023 TO 2025

Two held and grew. One faded: the enrolment deficit in the informal mining provinces had closed by 2025, almost entirely because of Tuv, and three explanations fit without the data being able to separate them. Four are no longer visible in the newer levels, mostly because company-funded schools and kindergartens arrived after the panel ended.

The second check was the background figures rather than the findings, and it covered all three papers.

Table of eleven background figures traced back to their sources across all three papers
ELEVEN BACKGROUND FIGURES TRACED BACK TO THEIR SOURCES, ACROSS ALL THREE PAPERS

None of it touched the regressions. All of it is in the setup, which is where a paper borrows numbers from other people's summaries rather than generating them. Four things were going on. Two sources with different definitions had been merged into one sentence, which is the hardest kind to catch because both halves check out on their own. Some numbers turned out to have no primary source at all and only exist in journalism that cites nothing. Some were true on the day they were written and had been overtaken since. And some were correct but quoted without the definition that gives them meaning: a maximum read as an average, or a guideline year that has changed.

That last category is the one I would not have found by re-reading my own work, because nothing about it looks wrong.

05 / What this doesn't prove

One project, and no one keeping score

Two separate things have limits here, and they are worth keeping apart. The research has limits, and the experiment has limits.

On the research: extending a paper does not make it stronger than its design. The first paper has two treated provinces per group and a borrowed treatment definition, so its estimates are associations, not causal effects, and nothing in the extension changes that. Gross enrolment ratios can also rise because the children of unregistered migrants enrol while the official population count stays where it is, so enrolment up is not the same thing as children better educated. The report says all of that in the open and lists every claim with its status in an appendix.

One of those limits has since moved. I wanted to test how accurate these tools actually are, so in September 2026 I put the same question to four of them: where do Mongolia's education numbers live below the province level. The answers ranged from the exact table names to a flat denial. Gemini said the data did not exist, repeated that word for word when the question was asked a second time, then reversed the moment it was handed one of the table IDs. Asked afterwards what had happened, it said it had answered from its own memory instead of searching.

Comparison of four AI tools answering the same Mongolia education data question
THE SAME QUESTION PUT TO FOUR TOOLS ON ONE DAY, WITH EVERY ANSWER CHECKED AGAINST THE LIVE SOURCES

They do exist. The statistics office publishes schools, students, teachers, graduates, kindergartens and new entrants for every soum and district in the country, most series running from the mid 2010s through 2025. Checking those answers against the live database also turned up something none of the four had been asked for: mid year population counts for the bags and khoroos underneath those districts, going back to 2001, which is the denominator that turns a count of pupils into a rate.

Tuv is a good place to see what that changes, because Tuv is the province section 04 turns on. In the paper it is one row. Underneath that row are 27 soums and 307 active mining licences, and neither is spread evenly.

Map of the 27 soums inside Tuv and their active mining licences
THE 27 SOUMS INSIDE TUV, AND EVERY ACTIVE MINING LICENCE IN THEM
School student numbers across the 27 soums inside Tuv from 2015 to 2025
THE SAME PROVINCE, BY SCHOOL STUDENT NUMBERS, 2015 TO 2025

Zaamar holds 107 of those licences and every one of them is gold, which is exactly the kind of place the original paper was built to compare, and at province level it is invisible. Across the province student numbers rose 45 percent, but that single figure covers a soum that lost pupils and another that more than quintupled, and Zaamar itself sits below the province average while the provincial centre sits far above it. A limitation of the original paper was that it compared twenty one provinces, because that was the only level available at the time. That is no longer the constraint. None of this repairs the original design, since a better dataset does not make a weak design strong in hindsight, and none of it is a finding yet. It is the clearest illustration of what this experiment was for: the step that ended the work five years ago, finding the data, is now the step that can be handed off.

On the experiment, the honest version is on the left below, and what would actually settle it is on the right.

Two columns listing what this does not prove and how you would actually find out
WHERE I WOULD PUSH BACK ON MYSELF, AND WHAT WOULD ACTUALLY TEST ANY OF IT

The first one matters most. A quarter of coursework against one working session is not a fair fight, because the paper already existed and I already knew the material. The real claim is smaller than the comparison makes it look: not that AI can produce research, but that it can finish research that had stalled.

06 / What I think this means

Three things I came out of this believing

The constraint was never the thinking. It was the mechanical middle: finding the tables, retyping them, aligning them, chasing down which source is right. That part is now cheap, and it was the exact part that made me stop the first time.

A place with thin data is where this helps most and where it is riskiest, at the same time. Mongolia has one statistics office, real gaps, and figures that disagree across sources and years. AI is very good at finding those disagreements and perfectly willing to smooth over them if nobody makes it show its work. Both of those were true within the same afternoon.

And the part that did not move is the part I would have expected to. Deciding what to care about, and deciding what I am willing to say, were mine at the start and mine at the end. Everything in between changed.

What I still do not know is whether any of it reaches anyone. The papers sat in a folder for five years. This one is public, which is a different thing from being read.

How I made it

Three papers

Handed over all three, with the datasets and the regressions.

NSO database

Pulled the provincial series again and ran them through to 2025.

Six more sources

World Bank, EITI, FAO, ILO, UN and PISA, reconciled against each other.

AI, at high effort

Extended the analysis, joined the three papers, and argued against my own conclusions.

A second pass

Reviewed the first pass adversarially, and caught a province it had mislabelled.

Me

Decided what to accept, what to qualify, and what to drop.