Blinded screening accuracy
| reference set | n | found |
|---|
| examiner applied it to anticipate (§102) | 40 | 97.5% |
| examiner applied it for obviousness (§103) | 40 | 92.5% |
| examiner never cited it | 80 | 18.8% |
The model never saw the reference's patent number, title, assignee or dates, so
it could not lean on anything it may have memorised. The control row is what
makes the other two mean anything: recall alone is trivially gained by flagging
everything, so the same screener was run over references the examiner did not
cite, drawn from the same corpus and passing the same priority-date gate.
Scope, stated in full
The denominator is examiner-applied references that are granted US patents
inside the 171,695-patent corpus. 73% of examiner citations point at pre-grant
publications and 6% at non-patent literature; both sit outside this corpus and
are excluded from numerator and denominator alike.
This is recall against the examiner, not against ground truth. An examiner's
own search runs 45 to 85% recall, so every reference Nightshift finds that the
examiner missed is scored here as a miss. The number is a floor, not an
estimate.