Catching serious errors shouldn't be a matter of luck.
Refine finds the technical errors in a document before anyone else does, performing several days' worth of expert reviewer work in about 30 minutes.
Empirical specification description
"We use the inverse of the propensity scores as weights."
For ATT, control units should be weighted by ωi / (1 - ωi), not 1 / ωi, where ωi is the propensity score of unit i. Suggestions: if the nonstandard weights are appropriate in this context, explain them when the weights are discussed.
Address case when x is negative in Proposition 2.
"Lemma 2. The optimal level y(x) satisfies..."
The proof of the lemma implicitly assumes the argument x is nonnegative, but the maintained assumptions do not require this. Suggestions: either add an argument to deal with the negative case, or make an assumption to rule out negative x.
Trusted by researchers at top institutions








Used by leading academic journals, including those of


Purpose-built for verification
As AI makes plausible-looking analysis cheap, checking it has become the bottleneck. Ask a chatbot or an agent to review the same paper twice and you get two different reviews. Refine is built so that the serious errors surface every time.
9 in 10
reviews won by Refine
head-to-head against GPT-6 Astra and Claude Fable 5.1 at maximum reasoning
24
central banks
where researchers check high-stakes analysis with Refine
10 of 10
fields won
Refine beat GPT-6 Astra and Claude Fable 5.1 in every field tested, from economics to medicine, biology, engineering and physics
Every argument, number and claim, stress-tested
Catch it first
Find the errors before referees, regulators or opposing experts do
A mistake found after publication, filing or sign-off costs far more than one found before. Refine reads your work the way a team of demanding experts would, and tells you what they would flag.
"spotted a number of errors"
I have found Refine to be very useful. It has spotted a number of errors, both obvious ones, and some that are fairly subtle, and it does a great job of checking for consistency between what is said in various parts of a paper.

Drew Fudenberg
Paul A. Samuelson Professor of Economics
MIT
Check every claim
Reasoning, numbers, evidence and code
Refine checks that each conclusion follows, that the text, tables and figures agree, and that every claim is supported, in any technical field, from economics and mathematics to biology, physics and law. Each issue is tied to its passage and ranked by how much it matters, and you can question any finding in chat.
"Refine is amazing!"
Refine is amazing! I've used it on projects ranging from the least to the most technical parts of philosophy, and in every case it's found errors I had missed, from small issues at the sentence or paragraph level to errors in keeping track of cases in longer proofs, all the way to mistakes in big arguments or how I cited the literature. The gain over publicly available tools is very significant. I plan to run every paper I write through it at least once before submitting.

Harvey Lederman
Professor of Philosophy
UT Austin
Check every draft
About 30 minutes, not weeks
Most documents come back in 20 to 40 minutes, so you can check every version of your work, not just the last one.
"would have taken an expert many hours to complete"
The depth of the reading performed by Refine was impressive, and resulted in the identification of subtle issues. This would have taken an expert many hours to complete, even when working with ChatGPT Pro.

Omer Tamuz
Professor of Economics and Mathematics
Caltech
How it works
Reliable verification from probabilistic parts
A single AI review is a sample: ask twice, get two different answers. Refine is engineered to catch the serious errors run after run.
Refine trains verifiers
Generating and verifying are different jobs. We seed real documents with subtle errors that frontier models miss, then train our verifiers until they catch them.
The state of the art in verification
Refine combines a proprietary harness over all top frontier models, deterministic checkers and purpose-trained classifiers. It uses 300 times the compute of a single AI call, so that all significant mistakes get caught.
Retuned with every release
We retune Refine for every major model release, so it improves as the field moves instead of being overtaken by it.
Advised by John Schulman
Our AI research is advised by John Schulman, co-founder of OpenAI and Chief Scientist at Thinking Machines.
Benchmark
Beats GPT-6 Astra and Claude Fable 5.1 at finding errors, in every field tested
Refine won 197 of 215 head-to-head reviews against GPT-6 Astra and Claude Fable 5.1.
Refine's reviews against reviews written by GPT-6 Astra and Claude Fable 5.1 at maximum reasoning, on the same 108 papers. Refine won the majority of matches in all ten fields: macroeconomics, econometrics, applied microeconomics, economic theory, medicine, biology, engineering, environmental science, computer science and statistics, and physics.
How we ran the benchmarkFig. 1 · Every match, one square
- Refine 5 won197
- Tie3
- GPT-6 Astra won10
- Claude Fable 5.1 won5
197of 215
matches won by Refine 5 (91.6%)
Privacy, by design
Your work stays yours: private, secure, and never used to train AI
Your content is completely secure
We protect your content with strong security controls and encrypted storage, following modern best practices. Learn more in our Privacy Policy.
Never used for training
Your papers are yours. Period. Nothing you upload ever becomes training data.
SOC 2 + ISO 27001
Certified in both. Our full security posture (subprocessors, encryption, incident response) is public on our Trust Center.
Questions