Developer tools · · 5 min read
Docling: turn a PDF into Markdown and JSON for an AI knowledge base
Convert one local PDF, inspect its headings and tables, and keep a traceable source record before adding it to an AI assistant.
By Sociologix Editorial

Prepare the source before asking an assistant to answer
A knowledge assistant can retrieve a paragraph and still answer incorrectly if the conversion dropped a condition or detached a table value from its column heading. This walkthrough prepares one PDF for inspection before it becomes searchable. The deliverable is a readable Markdown file, a structured JSON file and a small record connecting both to the original document.
This AI-assisted editorial guide uses official documentation checked October 8, 2026. Commands were checked against those sources; Docling and its models were not installed or benchmarked for this article. The review checklist below is a proposed workflow, not a claim about measured extraction accuracy.
1. Choose one document and create an isolated environment
The exact repository is docling-project/docling. It provides document conversion, including PDF structure analysis and OCR. Its code is MIT-licensed; model licenses must be checked separately. Use Python 3.10 or newer: the repository notes that Python 3.9 support ended in version 2.70.0. [1]
Start with an approved, non-sensitive PDF containing a heading, a paragraph and a table. Save it as sample.pdf in a new working directory. Keep the original unchanged. The following Windows PowerShell commands create an environment without changing the shell execution policy. The installation package is docling. [2]
On macOS or Linux, use python3 for environment creation and .venv/bin/python and .venv/bin/docling for the environment executables. Record the installed version before converting anything; documentation and installed command interfaces can differ.
python -m venv .venv
.\.venv\Scripts\python.exe -m pip install docling
.\.venv\Scripts\python.exe -m pip show docling
.\.venv\Scripts\docling.exe --help2. Export two views of the same PDF
The current CLI reference documents a convert subcommand, repeated --to options and an --output directory. If your installed help lists convert, run the command below. The expected outputs for sample.pdf are sample.md and sample.json under out. [3]
There is a documentation mismatch worth checking: the official quickstart still shows the older direct-source form, docling FILE. If installed help expects a source immediately and does not list convert, omit the word convert from the command below. Follow the help bundled with your installed version; do not treat a command-interface error as a broken PDF. [4]
Model files normally download on first use. That can require network access even though document processing uses local models. The advanced guide separately describes model prefetching for offline operation and explicit opt-in for remote model services. Leave remote services disabled for this local exercise. [5]
.\.venv\Scripts\docling.exe convert --from pdf --to md --to json --output out sample.pdf3. Inspect meaning, not just file creation
Open the PDF and Markdown side by side. A successful export is only the beginning of acceptance. Pick specific facts whose meaning depends on nearby text: an exception beneath a heading, a date range, a table value with units. A reviewer should be able to recover the same meaning from the export.
DoclingDocument represents text, tables, pictures, hierarchy and provenance, with layout information where available. Its JSON gives downstream code a richer representation than a flat text string. Keep it alongside Markdown rather than assuming a Markdown heading preserves every source location. [6]
- Check the first and last page for missing content and repeated headers.
- Compare one multi-column passage: sentences should follow the intended reading order.
- Check a table row against its row label, column label and unit.
- Check a scanned page for empty or garbled text. Hold failed documents out of the searchable collection.
4. Preserve the identity and permissions of the source
Create a small ingestion record containing a stable document ID, original filename, source version or revision date, installed Docling version, conversion settings, output paths and review status. Keep this record outside the extracted prose so it cannot be mistaken for an instruction from the document.
When you later split content into retrieval chunks, carry the document ID and permitted audience with each chunk. Keep available page or provenance references through that step, and test that the assistant can link back to the approved source. Conversion alone does not implement access control, retrieval or answer verification.
For a first acceptance check, write three questions with known answers and one the document cannot answer. Inspect both the retrieved passage and the final response. An unsupported answer should trigger a correction to the retrieval or response behavior, not an invented addition to the source.
5. Investigate failures before processing a folder
If a command is rejected, compare --help with the interface used above. If startup stalls, check whether model downloads are blocked before changing parsing settings. If a table loses meaning, retain the original and investigate that sample before importing the rest of the collection.
Rerun the same small sample after a dependency or model change and compare its outputs with your reviewed copy. Only expand to a batch when the relevant facts survive conversion. The useful milestone is a traceable, reviewable source for your assistant, not simply a directory full of extracted text.
Sources & further reading
Build an assistant around dependable company knowledge
Sociologix can help plan document ingestion, source-linked answers and permission-aware retrieval for an internal assistant. Bring one representative document and the questions your team needs to answer.
Talk to Sociologix