PDF to JSON Converter
Extract text, structure and layout coordinates from any PDF as clean, machine-readable .json — everything runs inside your browser, so your documents never leave your device.
PDF → JSON
4 Output Schemas
Layout Coordinates
100% Local
Why Convert PDF to JSON?
PDF and JSON are the two opposite poles of document processing. PDF is a human format — it is designed to look a specific way when printed or displayed, with glyphs positioned at exact coordinates, fonts embedded, and text reflowed to fit fixed pages. JSON is a machine format — it is a plain data structure of nested objects and arrays that a program can read, filter, transform, and feed into the next stage of a pipeline.
The gap between the two is enormous. Modern software — document management systems, search engines, RAG pipelines, LLM training crawlers, contract analysis tools, invoice processors, and BI dashboards — all consume structured data. They do not consume PDFs. Converting a PDF to JSON is the bridge that lets downstream software work with content that was previously locked inside a fixed visual layout.
Consider the concrete scenarios. A developer building a search feature needs the text of a PDF as searchable fields. A data scientist building a RAG pipeline needs each paragraph as a JSON object with metadata. A finance team automating invoice processing needs vendor name, amount, and line items extracted into a JSON API payload. A legal tech startup needs to index thousands of contracts by clause, page number, and position. None of these are possible with a raw PDF. All of them are trivial once the PDF is JSON.
The Web Tool Bazar PDF to JSON Converter does this conversion entirely inside your browser. There is no upload, no server round-trip, no account, and no queue. Your PDF — which might contain contracts, medical records, financial statements, or proprietary reports — never leaves your machine. When you close the tab, the session vanishes and only the JSON file you downloaded remains.
How PDF to JSON Conversion Actually Works
A PDF is not a data structure in disguise. It is a sequence of drawing instructions — "place this glyph at this coordinate," "draw a horizontal line here," "paint this image there." There is no concept of a paragraph, a heading, a sentence, or a table. Every structural element you perceive in a PDF is an optical illusion created by how the letters happen to line up. Recovering that structure requires four stages.
1. Text Extraction with Positions and Styles
Every text-based PDF stores a mapping between rendered glyphs and their Unicode characters, plus a transformation matrix that says exactly where each run of text sits on the page. The tool reads those mappings using pdf.js and captures not just the text but its x-position, y-position, width, and font size. It also inspects the font name to detect bold and italic variants (e.g., Helvetica-Bold, Times-Italic). This positional and stylistic data is the raw material for every schema.
2. Line Grouping
Text fragments that share a baseline are grouped into visual lines. The tool sorts all fragments by their y-coordinate, then walks down the page and collects any fragment whose y is within a small tolerance of the line's current baseline. The tolerance is tuned against the font size — larger text needs a larger tolerance because its baseline wobbles more. The result is a list of lines, each containing the text fragments that appear on that visual row, ordered left-to-right.
3. Paragraph Merging and Classification
Once lines are grouped, the tool merges them into paragraphs using a set of heuristics. Two consecutive lines belong to the same paragraph if the vertical gap between them is close to the median line height, if they share the same left margin (within tolerance), and if the previous line extends close to the right margin (suggesting the text was wrapped, not deliberately broken). Paragraphs are then classified: a paragraph whose font size is significantly above the document's median is marked as a heading; a paragraph that is short, does not end with a period, and has an unusual margin might be a title, caption, or list item.
4. Schema Serialisation
Finally, the classified content is serialised into one of four JSON schemas. Structured produces a hierarchical document with typed blocks — the right choice for republishing, editing, or downstream natural-language processing. Simple produces a flat array of pages, each with its full text — the smallest, most portable output. Layout produces an array of lines with x/y coordinates, font sizes, and styles — the right choice for tools that need to reconstruct the visual layout. Words produces a word-level stream — the right choice for NLP pipelines, tokenisers, and search indexers.
What You Get
The output is a valid RFC 8259 JSON file. It opens in every code editor, every JSON viewer, every programming language, and every data pipeline. You can pipe it into Python with a single json.load(), into Node with a single JSON.parse(), or into a jq one-liner for instant querying.
Understanding the Four JSON Schemas
Each schema is designed for a specific downstream use case. Choose the one that matches how you intend to consume the data.
Structured — Best for Republishing and Editing
Produces a hierarchical document with typed blocks. Every block has a type field ("heading" or "paragraph"), a level for headings, a text field with the full content, and optional style and page fields. This is the schema that most closely mirrors the document's logical structure.
{"document":{"pages":[{"blocks":[{"type":"heading","level":2,"text":"Introduction","page":1},{"type":"paragraph","text":"The report analyses...","page":1}]}]}}
Simple — Best for Search and Full-Text Indexing
Produces the smallest, most portable output: an array of pages, each with its full concatenated text. Ideal for feeding into a full-text search engine, a vector database, or a chunking pipeline for retrieval-augmented generation.
{"pages":[{"page":1,"text":"Introduction The report analyses..."},{"page":2,"text":"Section 2..."}]}
Layout — Best for Visual Reconstruction and PDF Tooling
Produces an array of lines, each with its exact x/y position, width, font size, and inline styling. Ideal for tools that need to reconstruct the visual layout — a custom PDF viewer, a redaction tool, or a document comparison engine.
{"pages":[{"page":1,"width":612,"height":792,"lines":[{"text":"Introduction","x":72,"y":720,"width":98,"size":18,"bold":true}]}]}
Words — Best for NLP and Tokenisation
Produces a word-level stream where every token has its own position, size, and style. Ideal for NLP pipelines, custom tokenisers, keyword extraction, and any task that needs sub-line granularity.
{"pages":[{"page":1,"words":[{"text":"Introduction","x":72,"y":720,"size":18,"bold":true,"line":0}]}]}
How to Convert PDF to JSON — Step by Step
1
Upload Your PDF File
Click "Select PDF File" or drag a PDF into the dashed drop zone. The tool loads the file locally, reads its page count, and renders thumbnail previews of every page.
2
Select the Pages You Want
Every page appears as a clickable thumbnail. Click to toggle a page on or off. Use "Select All" or "Select None" to pick everything or clear the selection. Only the selected pages will be included in the JSON.
3
Choose Your Output Schema
Pick the JSON shape that matches your downstream use case. Structured for republishing, Simple for search, Layout for visual reconstruction, Words for NLP. The live preview updates as you switch schemas.
4
Configure Optional Fields
Toggle PDF metadata, page numbers, font sizes, x/y coordinates, heading detection, bold/italic detection, and paragraph merging. Each option adds or removes fields from the JSON, so you can produce exactly the payload your pipeline needs.
5
Verify with the Live Preview
The live JSON preview shows the first 200 lines of the file that will be generated. Check that the structure, fields, and values look correct before downloading — every setting change instantly updates the preview.
6
Click "Convert to JSON"
The tool walks through the PDF page by page, groups text into lines, merges paragraphs, classifies blocks, and serialises the result into your chosen schema. A progress bar tracks each page. When it finishes, the download starts automatically and a summary reports the page count, block count, and file size.
7
Re-download Any Time
If you accidentally close the download prompt, the sidebar keeps the generated JSON in memory until you close the tab. Click "Re-download .json" to grab it again without re-converting.
What Works and What Doesn't — Being Honest
Every PDF to JSON tool, free or paid, makes tradeoffs. Understanding what converts cleanly helps you set realistic expectations.
What Converts Well
- Text-heavy documents — reports, articles, essays, manuals, terms and conditions
- Hierarchical structure — documents with clear headings, subheadings and body text
- Bold and italic runs — detected from the PDF font name and exposed as a
style field
- Multi-page documents — concatenated into a single JSON array, one element per page
- Reading order — lines are ordered top-to-bottom, left-to-right within each page
- Numeric and styled content — numbers stay as strings (JSON is character-based), but font size and styling are preserved
What Doesn't Convert Cleanly
- Scanned PDFs — pages that are photographs of paper contain no text at all. The tool warns you when it detects this, but cannot extract anything without OCR.
- Images, charts and vector graphics — not transferred. Only text is extracted.
- Tables as structured data — a table becomes a series of lines or blocks. Each cell is text, but the table's row/column structure is not explicitly encoded. Use the PDF to CSV or PDF to Excel tool for tabular data.
- Multi-column layouts — in Structured mode, a two-column magazine layout is re-flowed into a single stream. This is usually what you want for reading, but it changes the structure.
- Precise typography — the exact fonts embedded in the PDF are not reused. Font size and bold/italic are preserved, but the font family is not.
- Form fields — PDF form inputs are not transferred as JSON properties.
- Rotated text — text rotated 90° (common in newspaper-style tables) is extracted but its coordinates appear rotated as well.
If your PDF is a straightforward text document, the conversion will be clean and immediately useful. If it is a magazine layout with images and tables, expect to spend some time processing the JSON output downstream. Being upfront about these limits is more valuable than overpromising.
Real-World Use Cases
PDF to JSON conversion unlocks practical workflows across many roles and industries.
1. RAG and LLM Pipelines
Retrieval-Augmented Generation systems need document content as structured chunks. Converting PDFs to JSON gives you per-paragraph objects that can be embedded, indexed, and retrieved without re-parsing the original PDF every time.
2. Full-Text Search
Search engines like Elasticsearch, Solr, and Meilisearch index JSON documents natively. Converting a PDF to JSON turns it into indexable content, complete with page numbers and coordinates that can power result highlighting.
3. Document Management Systems
DMS platforms store metadata and content as structured data. PDF to JSON conversion is the standard way to ingest legacy PDFs into a modern DMS.
4. Contract Analysis and Legal Tech
Legal software needs clause-level access with page and position references. The Layout schema — with x/y coordinates — lets legal-tech tools highlight specific sentences on specific pages of a contract.
5. Invoice and Receipt Processing
Accounts payable automation needs vendor, amount, tax, and line-item data from invoices. Converting to JSON creates the structured payload that the accounting system's API can ingest.
6. Academic Research and NLP
Researchers building corpora need text and metadata as structured data. Converting PDFs to JSON gives them a clean input for tokenisers, NER models, and classification pipelines.
7. Content Migration and Republishing
Publishers migrating from PDF-based archives need the content as structured data before it can be rendered into HTML, EPUB, or a CMS. JSON is the intermediate format.
8. Automated Summarisation and Translation
Translation and summarisation services usually consume JSON as their input. Converting a PDF to JSON lets you feed the content directly into a translation API or an LLM summarisation endpoint.
Best Practices for PDF to JSON Conversion
- Check that the PDF is text-based. Open it in a reader and try to select text with your cursor. If you cannot highlight anything, the PDF is scanned and needs OCR before conversion.
- Choose the schema based on the consumer. RAG and search pipelines want Simple. Document editors want Structured. Visual tools want Layout. NLP needs Words. Do not pick the richest schema by default — pick the one that matches how the data will be used.
- Turn on "Merge wrapped lines" for prose. This is the single biggest readability improvement — it joins lines that belong to the same paragraph into a single block. Leave it off for tabular or list-heavy PDFs.
- Keep coordinates off unless you need them. The Layout and Words schemas add a lot of numeric fields, which significantly increases file size. If you only need the text, use Simple or Structured.
- Use minified output for production. During exploration, pretty-print with 2 spaces for readability. For production, switch to minified to reduce file size and payload over the wire.
- Verify with the live preview. The JSON preview shows the first 200 lines of the generated file. Check that headings are classified correctly, paragraphs are merged properly, and coordinates (if enabled) make sense.
- Do a schema-version note in your pipeline. If you feed this JSON into a downstream tool, note which schema you used. If you later change schemas, the downstream tool may break. Version your inputs explicitly.
- Keep the original PDF. Never delete the source. You may need it for reference, legal archival, or a future higher-fidelity conversion.
- Validate the JSON before using it. A one-line check like
python -m json.tool output.json confirms the file is well-formed before you invest in processing it.
Troubleshooting Common Issues
The JSON is empty or nearly empty.
This is the most common issue and almost always means the PDF is scanned. Open the PDF, try to select text with your cursor, and confirm it is text-based. If it is not, you need an OCR tool first.
Every paragraph is marked as a heading.
Heading detection is too aggressive for the document's font sizes. Turn "Detect headings" off, or switch to a simpler schema like Simple or Layout where heading classification is not used.
Paragraphs are split at every visual line.
"Merge wrapped lines" is probably turned off, or the PDF's line spacing is unusual. Turn on the option and check the preview again. If lines are still split, the PDF may store large vertical gaps between lines that happen to be part of the same paragraph — the merging is a heuristic, not a perfect reconstruction.
Words are joined together without spaces.
Some PDFs store text fragments without the expected whitespace, and the tool has to guess where spaces belong. This is rare, but if it happens, the output will have a few words joined. Fix them in your downstream code, or switch to a different PDF source if possible.
Coordinates are wrong or negative.
PDF coordinates are measured from the bottom-left corner of the page, with y increasing upward. Some PDFs use a different origin or a rotation that confuses the extraction. If you need exact pixel-perfect coordinates, you may need to pre-process the PDF to normalise its orientation.
Bold and italic runs are missing.
Make sure "Detect bold / italic" is checked. If styles are still missing, the PDF may not expose the font name in a way the tool can detect — some PDFs embed fonts with generic names that do not indicate weight or style.
The JSON file is very large.
You are using a rich schema (Layout or Words) at high verbosity. Switch to Simple or Structured, disable coordinates, and use minified formatting. A 100-page document at Layout + coordinates can easily produce a 20 MB JSON file; the same document in Simple mode is typically under 200 KB.
Non-Latin scripts (Arabic, Chinese, etc.) appear broken.
Text extraction of complex scripts depends on the PDF storing proper Unicode mappings. Older or poorly created PDFs may not include these. Try opening the PDF in a modern reader to confirm the text renders correctly there before assuming the tool is at fault.
The download didn't start automatically.
Click the "Re-download .json" button in the sidebar. The generated file is held in memory until you close the tab, so you can download it as many times as you need.
Frequently Asked Questions
What does the JSON output look like?
You choose the schema. Structured produces a hierarchical document with typed blocks. Simple produces a flat array of pages with their text. Layout produces an array of lines with x/y coordinates, font sizes, and styles. Words produces a word-level stream. Each is designed for a different downstream use case.
Can I convert a scanned PDF?
Not in this tool. Scanned PDFs are images, not text, and require Optical Character Recognition (OCR) to become editable. The tool detects pages with little text and warns you before conversion so you know what to expect.
Is my PDF uploaded to a server?
No. The entire conversion runs locally in your browser using pdf.js and native JavaScript. Your file never leaves your device. When you close the tab, everything is wiped from memory.
Does the tool preserve reading order?
Yes. Text fragments are grouped into lines by baseline and ordered top-to-bottom, left-to-right, which is the correct reading order for most languages. Multi-column layouts are handled by the same rule, which means two-column PDFs are read left column first, then right column.
Can I include layout coordinates?
Yes. Enable "Include XY coordinates" and each line or word gains x, y, and width fields. Coordinates are measured in PDF points (72 points per inch) from the bottom-left corner of the page, which is the PDF standard.
Does the tool support non-English PDFs?
Yes. The output is UTF-8 encoded JSON, so any Unicode character that pdf.js can extract is preserved. This includes Arabic, Chinese, Japanese, Korean, Cyrillic, and most other scripts.
How large can my PDF be?
The tool handles PDFs up to a few hundred pages comfortably. Very large files with rich schemas can take a few seconds per page to process. Everything runs on your device, so there is no upload delay.
Can I convert a password-protected PDF?
Encrypted PDFs must be unlocked before processing. Remove the password using your PDF reader first, then upload the unlocked file.
Can I convert multiple PDFs at once?
This tool converts one PDF at a time for clarity and quality control. To process several files, run them individually — each takes only a moment.
Is this tool really free?
Yes — completely free, no signup, no hidden fees, no daily limits, no watermarks. Use it as often as you need.
Final Thoughts
Turning a PDF into structured data should not require uploading it to a stranger's server or wrestling with desktop conversion software. The Web Tool Bazar PDF to JSON Converter delivers exactly what you need: positional text extraction, line grouping, paragraph merging, heading detection, and four production-ready JSON schemas — all running entirely in your browser, with no uploads, no accounts, and no watermarks.
Whether you are building a RAG pipeline, indexing a document archive, extracting invoice data, or feeding a downstream NLP model, this tool handles the job cleanly and privately. Bookmark it for the next time you need it — and explore our other PDF tools like the PDF to Word Converter, PDF to Excel Converter, PDF to CSV Converter, PDF to HTML Converter, PDF to JPG Converter, Merge PDF, and Split PDF, all built with the same privacy-first philosophy.