Apple-platform document intelligence that treats OCR and RAG as evidence pipelines, not as permission to turn uncertain text into confident facts.
The package demonstrates a confidentiality-safe subset of patterns used in production-style document workflows:
- PDFKit embedded-text fast path before OCR.
- Apple Vision accurate OCR fallback with configurable English/Russian recognition.
- Spatial line reconstruction from normalized Vision bounding boxes.
- Hard page, file-size and raster limits.
- Confidence and provenance retained as first-class data.
- Deterministic retrieval for downstream RAG context.
- Synthesis blocked when eligible sources contradict each other.
- Explicit
reviewRequiredoutcome when evidence is missing or too weak.
It contains no customer documents, private endpoints, employer code, credentials, trained models or proprietary extraction rules.
xcrun swift-format lint --recursive --configuration .swift-format Sources Tests Package.swift
xcrun swift test --parallel
xcrun swift build -c release
xcrun swift run evidence-pipeline-demoRequirements: Xcode 16.4+ or a compatible Swift 6 toolchain on macOS 14+. The library targets iOS 17+ and macOS 14+.
local PDF
-> file/page safety gates
-> PDFKit embedded-text fast path
OR PDF rasterization -> Apple Vision OCR
-> normalized lines + confidence + source provenance
-> deterministic fact extraction owned by the host app
-> GroundedContextBuilder
-> confidence gate
-> contradiction gate
-> deterministic query/fact ranking
-> cited context OR reviewRequired OR blockedByContradiction
-> optional downstream LLM under host-app policy
The package intentionally stops before an LLM call. A model may summarize the supplied context; it must not invent absent evidence or resolve contradictory official facts by itself.
Reads local PDFs with a fast, accurate text-layer path and an Apple Vision fallback. It preserves the chosen path per page and reports mean OCR confidence rather than flattening every result into a plain string.
Groups Vision observations into rows using normalized geometry, then orders row fragments left to right. The function is deterministic and unit-tested without requiring camera or document fixtures.
Builds a small, cited context for downstream RAG. Low-confidence facts are rejected, conflicting eligible values block synthesis, and deterministic tie-breaking keeps tests and audit logs stable.
let result = GroundedContextBuilder(minimumConfidence: 0.8).build(
query: "What is the permitted use?",
facts: extractedFacts
)
switch result.verdict {
case .grounded:
sendToModel(result.context, citations: result.citations)
case .reviewRequired, .blockedByContradiction:
routeToHumanReview(result.rejectedReasons)
}- Native Apple document APIs rather than a web-service wrapper.
- Swift 6 value semantics and explicit
Sendableboundaries. - Honest uncertainty and failure handling.
- Separation of OCR, deterministic extraction, retrieval and generative synthesis.
- Testable RAG guardrails that can fail closed.
- OCR is not guaranteed correct because a confidence value is high.
- Token overlap is not a replacement for semantic retrieval at large scale; it is a deterministic reviewable baseline.
- RAG does not make an LLM authoritative.
- The sample is not legal, cadastral, medical or financial advice.
- Host applications still own sandboxing, source authentication, PII policy, retention, schema validation and human review.
See ARCHITECTURE.md, THREAT-MODEL.md and VERIFICATION.md.
MIT.