What I Learned Running a Hundred-Page Empirical Study With AI

Author 
Coverage Type 

Last week I finished a hundred-page empirical report on four federal and two state broadband subsidy programs. It found that the programs mostly paid for service that was coming anyway, and that the ones with credible comparison groups bought a year or so of earlier service at a few thousand dollars per location per year. A project like that would once have taken me a year with a research assistant. It took about two weeks, and no research assistant. Anthropic’s Claude, running as an agent, assembled the data, wrote and ran the code, estimated the regressions, produced the tables and figures, and drafted text. I specified the questions, the designs, and the comparison groups, reviewed the results and the design choices at each step, and read some but not all of the code. OpenAI’s Codex then wrote its own code to re-estimate every family of regressions in the report from the project’s assembled data, and checked substantial parts of the data construction against the written methods. Of 563 numbers in the tables, counting repeats, 561 matched at the precision shown in the tables. The two that did not were rounding errors in converting two coefficients to percentages, and the replication caught them. It did not re-download the raw data or test the causal assumptions. The report says all of this in a note on methods, and I think a note like that should become standard.


What I Learned Running a Hundred-Page Empirical Study With AI