Last weekend’s chore was digitizing a massive pile of physical paperwork. I had accumulated and scanned a lot of documents: invoices, bank statements, tax papers, even kids’ homework, and completely random stuff. Because there were simply too many pages to sort through beforehand, I scanned them all together into multiple many-hundred-page PDFs.
When the dust settled, I had several large PDF files containing hundreds of randomly scrambled pages.
Sorting hundreds of pages manually—opening each one, identifying whether it’s page 2 of a multi-page tax filing, a utility invoice, or school homework, and dragging pages around in Acrobat—sounds like an exquisite form of torture. This is the exact kind of high-volume text classification problem modern LLMs chew through in seconds.
There was just one catch: personal paperwork contains sensitive private data—Social Security Numbers, bank account numbers, dates of birth, tax details, and personal signatures.
Keeping Sensitive Data Off the Internet
The first instinct many people have is to upload the PDFs directly to an online AI service or web interface and ask it to sort and split everything.
However, I didn’t feel comfortable sending my data to the Internet. Handing hundreds of pages of unredacted personal paperwork with Social Security Numbers and financial records to an external cloud endpoint was an immediate non-starter.
The Pattern: Decouple the Architect from the Worker
Instead of sending the paperwork to the model, I inverted the architecture. The entire document pipeline runs locally on Linux, while the cloud model is used strictly to build and refine the tooling without ever touching the files:
1. Gemini as the Systems Architect
I used Gemini inside my CLI strictly as a pair programmer. It never opened the PDFs, never extracted text, and never ingested OCR output. Its sole job was writing, profiling, and debugging a robust local organizer script. Whenever we hit an edge case (like multi-page tax forms or noisy scanned documents), Gemini analyzed diagnostic clusters without seeing a single piece of actual personal data.
2. Deterministic Fast Path
Standardized paperwork is surprisingly predictable. Tax documents (like W-2s or 1040s), recurring bank statements, and utility invoices often have consistent headers and sequential “Page X of Y” pagination. Using /usr/bin/pdftotext, Python extracted page headers, masked SSNs in-memory via regex, and matched over 85% of documents deterministically in under two seconds.
3. Local Gemma for Semantic Edge Cases
Small models (like 2B and 4B parameter models) often stumble when you ask them to run complex agent tool-calling loops. But they excel at focused, single-shot classification. For messy edge cases—handwritten notes, odd invoices, kids’ homework assignments, and unformatted documents—the script queried gemma4:e2b running locally via Ollama on localhost:11434. When Python acts as the driver, Gemma only has to answer simple classification questions, and not a single packet leaves the machine.
4. Lossless Assembly via pikepdf
Instead of re-rendering pages through an image pipeline (which destroys vector text, degrades scan DPI, and balloons file sizes), pikepdf performs lossless object manipulation. It preserves original embedded OCR text layers, scan DPI, and mixed page dimensions directly into clean, indexed sub-books.
Proving the Air Gap
Trusting a script is fine; verifying it with the Linux kernel is better.
To ensure that zero network traffic could possibly escape during execution, I launched the runner inside an isolated Linux network namespace with only the loopback interface brought up:
unshare -r -n bash -c "ip link set lo up && python3 organize_documents.py scan --no-ollama"
In this sandbox, the Linux kernel drops any non-loopback packet immediately. If any dependency attempted to dial out to the internet, it would instantly fail with Network is unreachable.
Finally, I audited the agent CLI transcript on disk. Because the CLI logs all tool calls in plaintext JSONL, a simple grep for document keywords and SSN regex patterns confirmed that not a single byte of document text or personal data was ever sent to Google’s servers:
grep -i "Bank Statement" ~/.gemini/antigravity-cli/brain/*/logs/transcript.jsonl
grep -E "[0-9]{3}-[0-9]{2}-[0-9]{4}" ~/.gemini/antigravity-cli/brain/*/logs/transcript.jsonl
Both queries returned zero results.
Final Thoughts
Hundreds of pages were organized into clean, categorized binders in under a minute without spending a dime on specialized document processing SaaS or exposing sensitive personal records.
We often assume that taking advantage of frontier AI models requires uploading our data to their servers. But frontier models are arguably at their best when used as software engineers—designing, debugging, and refining air-gapped local pipelines that you run and audit yourself.