Workshop · Tech note #04 · personal project

Pesquisa

An offline research engine, from arXiv to a typeset book

A local research environment in three layers: an engine that searches the open scientific databases and keeps a full-text library, editorial projects that become LaTeX PDFs, and a dashboard with versioned snapshots.

Illustrative schematic: arXiv, Crossref, Semantic Scholar and Unpaywall boxes converge on a “SQLite FTS5, OCR, Qualis” cylinder, which flows to “config.yaml build.py”, then “Tectonic LaTeX” and a PDF document; Ollama/CTranslate2 and versions boxes complete the flow.
Illustrative schematic made for this post. Not a screenshot.

The “read later” folder

Your paper collection is a folder called “read later”. Researchers pile up PDFs in folders, lose track of what they’ve read and depend on a connection and paid services to translate and search.

Pesquisa (Portuguese for “research”) is my local research environment. It builds an indexed library, searches the open scientific databases and takes the material all the way to a final LaTeX PDF, versioning every build. It has three layers.

Layer 1: the research engine

The engine searches four open databases in a single flow:

  • arXiv
  • Crossref
  • Semantic Scholar
  • Unpaywall
  • Local SQLite library with FTS5: full-text search over what you already have.
  • Downloading and organizing PDFs.
  • Optional OCR and offline translation.
  • Extractive Q&A over the documents.
  • Qualis rating from a CSV.
  • Export to Overleaf.
  • Offline AI with Ollama and CTranslate2.

Layer 2: editorial projects

Each book or guide is a project with a config.yaml and a build.py. The pipeline is always the same:

  1. Charts generated by the project.
  2. A typeset LaTeX PDF, via Tectonic.
  3. A JSON build report.

There are more than ten editorial projects in progress, on topics such as algebra, quantum barriers, electromagnetism, solid-state physics, SQUIDs, lithography and a guide to graphite.

Layer 3: the dashboard

Projects

A local dashboard to create and edit projects and trigger builds.

Versions

Every build produces a versioned snapshot and a report, and each project folder has its own git.

The result is reproducible: every build leaves a snapshot and a report behind.

Where it came from

Pesquisa started as an earlier, experimental batch, Criação de Estudos (roughly “study builder”). It wasn’t an application: it was the production setup for a technical e-book on graphite.

  • A YAML editorial plan for an e-book of about 300 pages, with research tasks, chapter-by-chapter writing, charts and equations.
  • Versions of about 100 and 200 pages, in HTML and LaTeX.
  • Python scripts that expanded the content and rendered the PDF.

That work grew into Pesquisa, with a search engine, a dashboard, versioning and a consolidated LaTeX pipeline. The graphite projects carry on inside it.

Tech sheet

  • Python 3.10+
  • SQLite FTS5
  • LaTeX (Tectonic)
  • Ollama
  • CTranslate2
  • arXiv · Crossref · Semantic Scholar · Unpaywall
Status
Active workspace with the engine, dashboard, tests and 10+ editorial projects
Network
Library, search, translation and AI work offline
Output
LaTeX PDF + JSON report + versioned snapshot

Do research, write technical books or want a library that works offline? Let’s talk.

Write to me