← All activity

EN

One of three AI-generated manuscripts passed peer review at a top ML conference. Let that settle for a moment.

Lu et al. (Nature, 2026) submitted AI-generated papers to ICLR workshops. One in three cleared the acceptance threshold. Separately, an automated reviewing system they tested reached 69% balanced accuracy — comparable to the variance you’d find across a panel of human reviewers.

Around the same time, Korinek (NBER WP34202, 2025) documented something that should reframe how organizations think about commissioning research: LLM agents are now running full literature reviews, econometric code, and end-to-end research workflows. Not as assistants. As operational components.

────

These two findings together are not a warning about AI replacing researchers. They are a signal about what rigorous applied research looks like in 2026 — and what it will cost.

The cost of producing credible evidence is falling. Not evenly, not without risk, not without new quality questions. But falling.

For organizations that commission evaluations — foundations, development banks, government agencies, impact investors — this changes several calculations at once.

Evaluation teams that integrate AI-augmented workflows will produce faster turnaround, broader literature coverage, and cleaner code pipelines. That is real capacity. The question is whether your procurement frameworks, your scope of work templates, and your quality review processes are designed for the world that existed before these tools, or the one that exists now.

────

Three things I’d flag for M&E directors and research commissioners:

  • The researcher+AI ensemble is now the unit of production. Evaluating a team’s capacity means understanding how they use these tools — not whether they use them.

  • “Rigorous” is not a fixed standard. The bar for what counts as a credible systematic review, a well-documented model, or a defensible causal claim is being renegotiated. Your quality criteria should be, too.

  • Transparency is the new methodology section. If a deliverable doesn’t document what AI did, how it was validated, and where human judgment was applied — that’s a gap in the audit trail, not a minor omission.

The infrastructure for evidence-based decision-making is being rebuilt. The organizations that engage with that process deliberately will be better positioned than those that wait for a standard to emerge.

#ImpactEvaluation #EvidenceBasedPolicy #AIResearch #DevelopmentFinance

On LinkedIn ↗