CorpusKit Studio

Getting started

Turn a PDF or EPUB into a private, searchable corpus — everything happens on your device.

The basics

1

Start a project

Tap New Project and choose a PDF or EPUB. Give it a name, then tap Create Project. CorpusKit Studio extracts the text and detects the chapters — it doesn't process anything yet, so you can fix the structure first.

2

Map the chapters

Open the Chapter Map tool. Review the detected chapters and set how important each one is. Importance weights the passages from that chapter when the corpus is searched.

3

Chunk & embed

In the Corpus Inspector, adjust the chunk size and overlap if you like, then run it. The text is split into passages and an on-device model turns each one into a vector. This is the one step that takes a little time — it runs entirely on your device, with no network.

4

Highlight what matters (optional)

Open Read Book, select any passage, and choose Highlight from the menu. Highlights boost those passages in search results. Pinch to zoom the page or text while you read.

5

Tune & evaluate (optional)

Use Query Expansion to map related wording (for example, “car” → “vehicle, automobile”), and Evaluation to write test questions and confirm the corpus answers them well — before you ship it.

6

Export

Open Export to package everything into a portable .corpus file. Open it in CorpusKit Reader to ask it questions and search it.

Good to know

  • Each tool shows a one-line reminder of what it does at the top of its panel — you never have to remember the workflow.
  • Opened a .corpus that someone shared? It carries the finished corpus but not the original document. To read the source pages and re-chunk, use Add Source PDF in the reader to attach your own copy.
  • Everything is computed on your device. Nothing is uploaded, and there is no account or tracking.