Skip to content

Open source

bloom-parser

A service that turns images, PDFs and spreadsheets into one structured document.

I built it and published it as open source.

Code

Role
Sole author
Context
Personal project, open source
When
2026

What I did

  • Routed 11 input formats through 4 adapters into one document model.
  • Made OCR optional at build time: the default binary needs no native libraries, and a build tag adds Tesseract.

Impact

Public MIT-licensed code; a new format is an adapter plus a detection rule, and the pipeline, OCR and exporters don't change.

bloom-parser pipeline A request is validated and its format detected, then routed to one adapter per format: image (with an optional OCR engine), PDF, spreadsheet or text. Every adapter produces the same Document, which can be exported to CSV or XLSX, or published to Power BI. requestvalidate ยท detectimagepdfxlsxtextDocumentCSV / XLSXPower BIOCR enginebehind an interface

When things go wrong

  1. When the OCR engine isn't installed,

    OCR requests get a clear per-page error and every other path keeps working.

Built with

Go, gRPC, Python, Tesseract, Vue