Language models are genuinely good at the first pass: proposing candidate extractions, clustering near-duplicate records, drafting a screening rationale, and flagging where two papers appear to disagree. They are unreliable as the final authority on any of it, and confidently so, which is the part that costs time further down the line.
So the model proposes and a person disposes, with the system recording which is which. Every field carries whether it was extracted automatically or confirmed by a reviewer, and only the confirmed set feeds analysis. That also produces the data needed to tell whether extraction quality is improving, rather than an impression that it is.