How to extract invoice data from PDFs into structured JSON

Invoice PDFs rarely follow one layout, so generic text extraction breaks on multi-column tables, wrapped cells, and scanned pages. This guide walks through the four API calls that convert a raw invoice PDF into typed JSON, plus the post-processing code that maps the response onto a clean invoice schema your backend can use right away.
Invoice PDFs are structurally unpredictable. Vendor A ships a five-column line-item table, vendor B embeds the same information in a paragraph block, and half the scanned copies in a legacy archive have no text layer at all. Generic text extraction reads characters in the order they were written to the file rather than the reading order a human sees, so field values drift with every layout variation.
The Foxit PDF Structural Extraction API, also known as the PDF Structural Analysis API, addresses this directly. You submit a PDF invoice and receive a typed, hierarchical JSON document with named element types, bounding regions, and an addressable table cell grid. This guide covers the four REST calls that turn a raw PDF into a StructureInfo.json file, the Python post-processing that maps that output onto a clean invoice schema, and the edge cases your pipeline will hit on real vendor documents. Every response shape and field name below comes from a live run against the API, not from the reference docs alone.
Why invoice PDFs break generic parsers
Three structural problems cause most invoice parsing failures, and OCR accuracy is only one of them.
The first is layout variance across vendors. A PDF’s text layer records characters in the order they were drawn, which often follows vector rendering order rather than left-to-right, top-to-bottom reading order. Extract raw text from a five-column line-item table and you frequently get interleaved fragments, where description text from one column mixes with unit prices from another because the writer rendered all rows of one column before moving to the next. No string-parsing logic reliably recovers column boundaries from that flattened sequence.
The second is merged and multi-row cells. Line-item tables routinely span cells across rows for items with multi-line descriptions. Text extraction collapses those cell boundaries into a flat string and drops the row-to-total relationship an accounting system needs.
The third is rasterized scans with no text layer. A PDF created by scanning a paper invoice contains only an embedded image, so anything that reads the text layer alone comes back empty. That kind of file has image data but no searchable text at all. Tools built for scanned input bundle an OCR step rather than skipping it, and any pipeline you build has to do the same before extraction can happen.
Structure-aware extraction addresses all three by classifying document regions into typed elements before exposing their content.
Invoice fields to target before touching the API
Define the target schema before writing code. A concrete target tells you which elements to read from StructureInfo.json and which to skip, which saves iteration time on every invoice you process.
A workable invoice schema covers three groups:
- Header fields, including vendor name, invoice number, invoice date, due date, and payment terms
- Line items, a repeating array of description, quantity, unit price, and line total
- Footer totals, including subtotal, tax amount, and total amount due
{
"vendor_name": "",
"invoice_number": "",
"invoice_date": "",
"due_date": "",
"payment_terms": "",
"line_items": [
{
"description": "",
"quantity": "",
"unit_price": "",
"line_total": ""
}
],
"subtotal": "",
"tax": "",
"total_due": ""
} Keep every value a string at extraction time. Type conversion, currency parsing, and date normalization belong downstream, after validation, where a bad value can be rejected with context instead of raising inside the parser.
The sample invoice used throughout this guide is invoice_full_test.pdf, so you can run every call below against the same document.
The input document. Notice that the Subtotal, Tax Rate, Tax Amount, and Total Due labels sit in the second-to-last column rather than the first. That detail determines how the post-processing code has to find them.
Prerequisites
You need the following before the first API call:
- Python 3.9 or newer and pip
- A virtual environment via venv, so the dependency below stays isolated
- The requests library for HTTP calls
- A code editor such as VS Code with the Python extension
- A free Foxit developer account
Scaffold the workspace in one shot:
mkdir invoice-extraction && cd invoice-extraction && python3 -m venv .venv && source .venv/bin/activate && pip install requests Then download the sample invoice into that folder:
curl -L -o invoice_full_test.pdf https://github.com/lucienchemaly/foxit-demo-templates/raw/main/invoice_full_test.pdf Foxit API authentication and setup
Signing up activates a free Developer plan that includes 500 credits per year with no credit card required. A structural extraction call costs one credit, while the upload, polling, and download calls are not billed, so a full run of the workflow below costs a single credit.
The account creation screen. The free Developer plan is enough to work through this entire guide.
Foxit authenticates PDF Services requests with a client ID and client secret passed as HTTP headers, so there is no OAuth token exchange to implement. Both values come from the default application created in your Developer Portal dashboard, alongside the base URL your calls need.
The credentials panel. Copy the Client ID and Client Secret into environment variables rather than pasting them into source files.
Export them into your shell so no credential is ever committed:
export FOXIT_CLIENT_ID="your_client_id"
export FOXIT_CLIENT_SECRET="your_client_secret" The structural extraction reference page carries a Test Request button that fires live calls straight from the browser, which is the quickest way to confirm your credentials work before writing any Python. The endpoint is currently labelled Trial in the reference, so expect its surface to evolve.
Pin your parser to the version field inside the analyzeResult response. The current schema ships as 1.0.7, and pinning prevents silent breakage if that changes.
The four-step PDF to JSON invoice extraction workflow
The API is asynchronous. You upload a document, start a task, poll until the task completes, then download the result.
All four paths sit under https://na1.fusion.foxit.com/pdf-services. Calling them without that prefix returns 404.
The path prefix matters more than it looks. The four endpoints live under /pdf-services/api/..., and requesting a bare /documents/{id}/download returns 404 rather than a helpful error.
import io
import json
import os
import time
import zipfile
import requests
BASE_URL = "https://na1.fusion.foxit.com/pdf-services"
HEADERS = {
"client_id": os.environ["FOXIT_CLIENT_ID"],
"client_secret": os.environ["FOXIT_CLIENT_SECRET"],
}
POLL_SECONDS = 2
POLL_TIMEOUT = 120
def extract_structure(pdf_path: str) -> dict:
# Step 1: upload the PDF (multipart/form-data, 100 MB maximum)
with open(pdf_path, "rb") as handle:
upload = requests.post(
f"{BASE_URL}/api/documents/upload",
headers=HEADERS,
files={"file": (os.path.basename(pdf_path), handle, "application/pdf")},
)
upload.raise_for_status()
document_id = upload.json()["documentId"]
# Step 2: start the structural extraction task
started = requests.post(
f"{BASE_URL}/api/documents/pdf-structural-extract",
headers=HEADERS,
json={"documentId": document_id},
)
started.raise_for_status()
task_id = started.json()["taskId"]
# Step 3: poll until COMPLETED, bounded, and handle FAILED
deadline = time.monotonic() + POLL_TIMEOUT
while True:
response = requests.get(f"{BASE_URL}/api/tasks/{task_id}", headers=HEADERS)
response.raise_for_status()
task = response.json()
if task["status"] == "COMPLETED":
break
if task["status"] == "FAILED":
raise RuntimeError(f"extraction task {task_id} FAILED: {task}")
if time.monotonic() > deadline:
raise TimeoutError(f"task {task_id} stuck at {task['status']}")
time.sleep(POLL_SECONDS)
# Step 4: download the result ZIP and read StructureInfo.json
result = requests.get(
f"{BASE_URL}/api/documents/{task['resultDocumentId']}/download",
headers=HEADERS,
)
result.raise_for_status()
with zipfile.ZipFile(io.BytesIO(result.content)) as archive:
return json.loads(archive.read("StructureInfo.json")) This code:
- Reads both credentials from the environment.
- Uploads the PDF as multipart form data and captures the returned
documentId. - Hands that id to the extraction endpoint to receive a
taskId. - Polls the task endpoint every two seconds, bounded by a deadline, checking explicitly for
FAILEDso a rejected document raises instead of spinning forever. - Once the status reads
COMPLETED, downloads the ZIP archive using the task’sresultDocumentIdand readsStructureInfo.jsonout of it in memory.
A real run against the sample invoice. The task reports IN_PROGRESS at 20 percent before reaching COMPLETED, and the final object carries every field from the schema defined earlier.
How to map raw output to a clean invoice schema
Choosing an invoice data extraction API is only half the work. The other half is mapping whatever it returns onto a schema your backend already understands, and that mapping is where the shape of the response starts to matter.
StructureInfo.json wraps everything in an analyzeResult object with four top-level keys, version, pages, info, and elements. The elements array is where the work happens. Each element carries a type drawn from twelve values, including paragraph, table, title, image, form, and formula, along with its bounding region and content. The same API can also extract embedded images as standalone image files, in addition to returning them as typed image elements within StructureInfo.json.
Two details in that structure cause most of the bugs in a first implementation, and neither is obvious from the field names.
The actual response shape from a live extraction. A table’s cells are nested at content.body.cells, cell text sits at paragraph.content.text, and region.boundingBox is an eight-number polygon rather than an x, y, width, height rectangle.
A table element does not expose a top-level cells array. Its grid is nested at content.body.cells, where each cell carries rowIndex, columnIndex, and a paragraph object. Cell text then sits one level deeper still, at paragraph.content.text, because content is an object rather than a string. Reaching for cell["paragraph"]["content"] returns a dict, not the text you want.
Blank cells are the second detail. When a vendor leaves a cell empty, the API still returns the cell with its indices and a paragraph object, but that paragraph has no content key at all. An unguarded read raises KeyError partway through a document that looked fine in testing.
def element_text(element: dict) -> str:
"""Return an element's text, or an empty string when it carries none."""
text = element.get("content", {}).get("text", "")
return " ".join(text.split())
def cell_text(cell: dict) -> str:
"""Return a table cell's text, or an empty string when the cell is blank."""
return element_text(cell.get("paragraph", {})) In this code, element_text reads the nested content.text value and normalizes its whitespace, which matters because cell text can contain a literal \r\n where a label wraps across two lines. Using " ".join(text.split()) collapses those into single spaces, whereas .strip() leaves a mid-string newline untouched. cell_text then reuses that helper for table cells, returning an empty string for a blank cell instead of raising.
With the accessors in place, build a grid and read it row by row.
import re
FOOTER_LABELS = ("subtotal", "tax rate", "tax amount", "tax", "total due", "total")
HEADER_PATTERNS = {
"vendor_name": r"bill to:\s*(.+)",
"invoice_number": r"invoice number:\s*(.+)",
"invoice_date": r"invoice date:\s*(.+)",
"due_date": r"due date:\s*(.+)",
"payment_terms": r"payment is due within (.+?) of",
}
def parse_invoice(structure_info: dict) -> dict:
elements = structure_info["analyzeResult"]["elements"]
invoice = {key: "" for key in HEADER_PATTERNS}
invoice.update(line_items=[], subtotal="", tax="", total_due="")
# Header fields come from paragraph elements above the table
for element in elements:
if element["type"] != "paragraph":
continue
text = element_text(element)
for field, pattern in HEADER_PATTERNS.items():
match = re.search(pattern, text, re.IGNORECASE)
if match and not invoice[field]:
invoice[field] = match.group(1).strip()
tables = [element for element in elements if element["type"] == "table"]
if not tables:
return invoice
grid: dict = {}
for cell in tables[0]["content"]["body"]["cells"]:
grid.setdefault(cell["rowIndex"], {})[cell["columnIndex"]] = cell_text(cell)
# Resolve columns from the header row instead of assuming positions
columns = {name.lower(): index for index, name in grid.get(0, {}).items() if name}
def column_for(*candidates, default):
for candidate in candidates:
for name, index in columns.items():
if candidate in name:
return index
return default
description_col = column_for("description", "item", default=1)
quantity_col = column_for("qty", "quantity", default=2)
unit_price_col = column_for("unit price", default=3)
line_total_col = column_for("total", "amount", default=4)
for row_index in sorted(index for index in grid if index > 0):
row = grid[row_index]
label = next(
(value.lower().rstrip(":").strip() for value in row.values()
if value.lower().rstrip(":").strip() in FOOTER_LABELS),
None,
)
if label:
value = row[max(row)]
if label == "subtotal":
invoice["subtotal"] = value
elif label == "tax amount":
invoice["tax"] = value
elif label in ("total due", "total"):
invoice["total_due"] = value
continue
if row.get(description_col):
invoice["line_items"].append({
"description": row.get(description_col, ""),
"quantity": row.get(quantity_col, ""),
"unit_price": row.get(unit_price_col, ""),
"line_total": row.get(line_total_col, ""),
})
return invoice In this code, you first walk the paragraph elements and pull header fields out with labelled regular expressions, which works because Foxit exposes each header line as its own element with its reading order preserved. You then flatten the table into a {rowIndex: {columnIndex: text}} grid and resolve column positions from the header row by name, so a vendor who adds a leading row-number column does not shift every field by one. Each subsequent row is classified before it is read, so that if any cell in the row matches a known footer label the row is treated as a total and its value taken from the last populated column, and otherwise the row becomes a line item. Scanning the whole row for the label is the part that matters, because footer labels do not sit in the first column.
Footer totals living inside the line-item table is convenient rather than awkward, since one pass over the cell grid covers line items and totals together. Form elements do not appear on a typical invoice, so there is no need to look for them.
Before and after
The two blocks below show the same data on either side of that mapping. First, one real cell exactly as the API returns it, taken verbatim from the run above:
{
"paragraph": {
"type": "paragraph",
"content": {
"text": "API Integration\r\nConsulting"
},
"region": {
"page": 1,
"boundingBox": [171, 300, 258, 300, 258, 328, 171, 328]
},
"id": "paragraph12",
"paragraphOrder": 12
},
"rowSpan": 1,
"columnSpan": 1,
"rowIndex": 1,
"columnIndex": 1,
"region": {
"page": 1,
"boundingBox": [171, 300, 258, 300, 258, 328, 171, 328]
},
"score": 0.8555269837379456
} That fragment is the shape of every cell you will handle. The text is nested at paragraph.content.text rather than sitting directly on the cell. The position arrives as an eight-number boundingBox polygon on both the cell and its paragraph, not as a rectangle. And the value itself contains a literal \r\n where the description wrapped onto a second line in the source table, which is the case .strip() silently fails to clean.
Second, the complete object parse_invoice returns for the whole document, which is what your backend actually consumes:
{
"vendor_name": "Acme Corporation",
"invoice_number": "INV-2025-0042",
"invoice_date": "07/15/2025",
"due_date": "08/14/2025",
"payment_terms": "30 days",
"line_items": [
{
"description": "API Integration Consulting",
"quantity": "8",
"unit_price": "$ 195.00",
"line_total": "$1,560.00"
},
{
"description": "Document Automation Setup",
"quantity": "1",
"unit_price": "$ 750.00",
"line_total": "$ 750.00"
}
],
"subtotal": "$2,310.00",
"tax": "$ 184.80",
"total_due": "$2,494.80"
} The wrapped description has become the single clean string "API Integration Consulting", the header fields have been lifted out of the paragraph elements above the table, and the four footer rows have been separated from the two genuine line items. That output is backend-ready, so you can write it straight to a database, push it to a reporting pipeline, or validate it against an accounts-payable schema without further parsing.
Common mistakes
- Dropping the path prefix. All four endpoints sit under
/pdf-services/api/.... A bare/documents/{id}/downloadreturns 404. - Reading
cellsoff the table element. The grid is nested atcontent.body.cells. A top-levelcellslookup raisesKeyError. - Treating
paragraph.contentas a string. It is an object, so the text is atparagraph.content.text. - Assuming footer labels are in column 0. On real invoices they commonly sit in the second-to-last column, with the value beside them.
- Forgetting blank cells. A blank cell keeps its
paragraphobject but carries nocontentkey. - Polling without a bound. A
FAILEDtask never becomesCOMPLETED, so an unboundedwhile Trueloop hangs. - Trusting
info.basicInfo.elementCounts. It can disagree with the length of theelementsarray, so size loops from the array itself. - Using
.strip()to clean cell text. A wrapped label contains a mid-string\r\nthat.strip()leaves in place. Use" ".join(text.split()).
Invoice parsing FAQ
What is invoice data extraction from PDF?
Invoice data extraction from PDF is the programmatic conversion of semi-structured PDF invoice content into a schema-typed data object. The source can be a digital-native PDF with an embedded text layer or a scanned image-only PDF that needs OCR first. The output is a structured record, typically JSON, with typed fields for header values, line items, and totals that downstream systems consume without manual parsing.
How do I extract line items from a PDF invoice?
Line items live in table elements inside StructureInfo.json. Read the grid from content.body.cells, where every cell carries a rowIndex and columnIndex, build a {rowIndex: {columnIndex: text}} dictionary, resolve the column positions from the header row, then read each data row in column order. Classify rows before reading them so footer totals are not appended as line items.
Can an extraction API handle scanned PDF invoices?
The PDF Structural Extraction API expects a PDF that already has a text layer. Tested against an image-only PDF, it returns no text and an empty table rather than an error. To handle scans, run the document through Foxit’s OCR endpoint first (POST /pdf-services/api/documents/analyze/pdf-ocr with outputFormat set to PDF), then pass the OCR output through the four-step workflow. That makes five calls rather than four, and the OCR step is billed as its own credit.
What does the JSON output from PDF invoice extraction look like?
The downloaded ZIP unpacks to StructureInfo.json, holding an analyzeResult object with version, pages, info, and elements keys. The elements array carries every classified region, each with a type drawn from twelve values. A table element nests its grid at content.body.cells, and each cell’s text sits at paragraph.content.text.
Why build a grid instead of iterating the cells array directly?
A grid keyed by row and column decouples your field mapping from the order the API happens to return cells in, and it lets you address a specific position directly, which is what the footer-label check needs. It also makes missing cells visible as absent keys rather than as silently shifted values.
What element types does the API classify?
The API classifies regions into twelve element types, including paragraph, table, title, image, form, and formula. On a typical invoice, paragraph elements carry the header fields while a single table element carries both line items and footer totals, so filtering by type lets you target only what your schema needs.
How should I handle invoices where footer totals sit outside the line-item table?
Some layouts place totals in a separate table element or in standalone paragraph elements below the main table. Keep the row-classification logic in its own function so you can apply it to a second table element, then fall back to scanning paragraph elements for currency-formatted strings next to known label text.
Wrapping up
The four-step workflow of upload, extract, poll, and download produces a typed, hierarchical JSON document from any digital-native invoice PDF. Post-processing the elements array through a row and column grid turns that into a clean invoice object your backend can consume directly. The same four calls extend to purchase orders, receipts, and any other tabular financial document, with only the mapping logic changing to match the target schema.
The Foxit PDF Structural Extraction API is part of the broader Foxit PDF Services API, which covers conversion, compression, OCR, and other document operations under the same credential set. The infrastructure is SOC 2 Type II certified, with GDPR-supporting features and HIPAA-aligned controls including BAA availability, which matters for teams processing financial documents under compliance review.
The complete script from this guide is available as extract_invoice.py if you want to run it before adapting it.
Create your free developer account and start turning invoice PDFs into structured JSON today. No credit card required, and the 500 credits on the free Developer plan are enough to build and test something real.
How Foxit compares to Google Document AI for document data extraction

Seven Google Document AI processors are being retired in June 2026, pushing teams to look at alternatives. This comparison walks through integration setup, extraction architecture, output schema, and data residency for Foxit’s PDF Structural Extraction API against Google Document AI, so you can decide with real implementation detail instead of a feature list.
Seven Google Document AI processors hit end-of-life on 30 June 2026. Google is retiring the Enterprise Document OCR, Expense, Custom classifier, Custom splitter, Invoice, Pay slip, and Bank statement parsers, and processor versions follow a rolling schedule where each version is deprecated six months after a newer one ships.
For many teams, the deprecation notice does more than prompt a migration ticket. It opens a broader question about whether Google Document AI is still the right foundation for the extraction stack, and what the credible Google Document AI alternatives actually look like once you compare them on implementation detail rather than feature lists.
This article gives you the technical specifics to make that call, comparing the Foxit PDF Structural Extraction API and Google Document AI across five dimensions that drive the real build-vs-switch decision, covering integration overhead, extraction architecture, output schema, document and language coverage, and data residency.
Five axes for evaluating Google Document AI alternatives
Any extraction API comparison lives or dies on the criteria it uses. These five dimensions cover what a production engineering team actually cares about, going well beyond a proof-of-concept benchmark.
Integration complexity measures how many external dependencies you must provision before your first call returns data. A tool that requires a GCP project, service account, IAM role grants, and billing enablement adds meaningful friction before a single byte of document gets processed. For teams with CI/CD pipelines and strict access-control policies, every new cloud dependency is a potential blocker.
Extraction architecture covers how the API reads a document internally, including what happens when a file mixes scanned pages, machine-typed text, and embedded tables in the same document.
Output schema determines how much post-processing your downstream systems require. A flat token list forces you to reconstruct document structure yourself, while a pre-labeled semantic taxonomy reduces that burden before the data reaches your RAG pipeline or BI dashboard.
Document and language coverage sets the practical ceiling on what you can run through the API in production. Language breadth matters especially for multi-region workloads processing invoices or contracts in non-Latin scripts.
Data residency encompasses where documents travel during processing, how long they remain on third-party infrastructure, and what audit evidence you can produce for compliance reviews. For regulated industries, this dimension often decides the question before the others are evaluated.
Integration setup and authentication overhead
Getting to a first call on Google Document AI requires a GCP project, a service account with an IAM role assignment (at minimum roles/documentai.apiUser), billing enablement on the project, and an environment variable pointing to a downloaded service account JSON key. Teams outside the GCP ecosystem absorb all of that as onboarding cost before any extraction runs.
Foxit’s path is shorter. Create a free developer account at the Foxit Developer Portal, retrieve your client_id and client_secret from the default application, and attach them as two HTTP headers on every request. The entire setup takes minutes and requires nothing from GCP.
Prerequisites
To run the code below you need Python 3.8+, the requests library installed into an isolated virtual environment with pip, an editor such as VS Code with the Python extension (PyCharm or Sublime Text work equally well), and a free Foxit developer account from app.developer-api.foxit.com/sign-up to supply the two credential values. Scaffold the workspace in one shot:
mkdir foxit-extract && cd foxit-extract && python3 -m venv .venv && source .venv/bin/activate && pip install requests Extraction then follows a four-step asynchronous workflow, annotated at each step:
export FOXIT_CLIENT_ID="your_client_id"
export FOXIT_CLIENT_SECRET="your_client_secret" Extraction then follows a four-step asynchronous workflow, annotated at each step:
import os
import requests
import time
BASE_URL = "https://na1.fusion.foxit.com/pdf-services/api"
HEADERS = {
"client_id": os.environ["FOXIT_CLIENT_ID"], # lowercase snake_case, not Authorization: Bearer
"client_secret": os.environ["FOXIT_CLIENT_SECRET"]
}
# Step 1: Upload the document (multipart/form-data, field name "file", max 100 MB)
with open("contract.pdf", "rb") as f:
upload_resp = requests.post(
f"{BASE_URL}/documents/upload",
headers=HEADERS,
files={"file": f}
)
document_id = upload_resp.json()["documentId"]
# Step 2: Start structural extraction
extract_resp = requests.post(
f"{BASE_URL}/documents/pdf-structural-extract",
headers=HEADERS,
json={"documentId": document_id}
)
task_id = extract_resp.json()["taskId"]
# Step 3: Poll every 2 seconds until COMPLETED (statuses are uppercase)
result_doc_id = None
for _ in range(60):
status_resp = requests.get(
f"{BASE_URL}/tasks/{task_id}",
headers=HEADERS
)
payload = status_resp.json()
if payload["status"] == "COMPLETED":
result_doc_id = payload["resultDocumentId"]
break
if payload["status"] == "FAILED":
raise RuntimeError(f"Extraction failed: {payload}")
time.sleep(2)
if result_doc_id is None:
raise TimeoutError("Extraction did not complete within 120 seconds")
# Step 4: Download the ZIP archive containing StructureInfo.json
result_resp = requests.get(
f"{BASE_URL}/documents/{result_doc_id}/download",
headers=HEADERS
)
with open("extraction_result.zip", "wb") as out:
out.write(result_resp.content) Authentication uses lowercase snake_case header names on every call. The upload endpoint accepts multipart/form-data with the PDF file under the field name file, with a 100 MB size limit per document.
The four calls against the live API. Note the 202 on the extract call and the COMPLETED status before any download is attempted.
The response shape from that run. Every element sits under analyzeResult, and text elements carry region.boundingBox as an eight-number polygon rather than a four-number rectangle. A table element is the exception, since its region comes back empty and its geometry sits in a regions array instead.
The table below puts both platforms side by side on the integration and output dimensions:
| Dimension | Google Document AI | Foxit PDF Structural Extraction API |
|---|---|---|
| Account setup | GCP project, service account, IAM role, billing | Free developer account, no credit card |
| Authentication | Service account JSON key via GOOGLE_APPLICATION_CREDENTIALS | client_id and client_secret as HTTP headers |
| Call pattern | Synchronous or async depending on processor | Four-step async (upload, extract, poll, download) |
| Output format | Document proto (blocks, paragraphs, tokens) | StructureInfo.json with 12 semantic element types |
| Cloud dependency | GCP-native; Vertex AI integration available | Cloud-agnostic, any stack |
| Language coverage | Varies by processor | 200+ languages via dedicated OCR layer |
Common mistakes and troubleshooting
Four failure modes account for most of the time lost on a first integration. Only the first one reports itself clearly; the other three surface as exceptions in your own code rather than as API errors.
- Sending an
Authorization: Bearerheader : PDF Services authenticates with two separate lowercase headers,client_idandclient_secret. There is no token exchange step and no OAuth flow to implement. This is the one mistake the API names outright, returning HTTP 400 with{"allow": false, "reason": "Missing credentials: provide both 'client_id' and 'client_secret' headers."}. - Treating the extract call as synchronous :
pdf-structural-extractreturns HTTP 202 with ataskId, never the result. Read the status fromGET /tasks/{taskId}, which also returns aprogresspercentage, and note that the values are uppercase (PENDING,IN_PROGRESS,COMPLETED,FAILED). Comparing against lowercase strings produces a poll loop that never exits, which is why the sample above also breaks out onFAILEDand caps its attempts. - Reading
region.boundingBoxon every element : text elements such astitle,head, andparagraphcarryregionas{page, boundingBox}, but atableelement returns an emptyregionand puts its geometry in aregionsarray instead. A loop that assumes one shape raisesKeyErroron the first document containing a table. - Indexing
elementsat the JSON root : every result nests underanalyzeResult, sodata["elements"]raisesKeyErrorwhiledata["analyzeResult"]["elements"]works. The same applies to uploads above the 100 MB per-document limit, which fail at the upload step rather than during extraction.
Extraction architecture, output format, and document coverage
Google Document AI’s processing model is processor-centric. You select a processor type (Invoice Parser, Form Parser, Document OCR), and the service returns a Document proto containing position-anchored blocks, paragraphs, and tokens. Semantic meaning depends on which processor you deployed, so a Form Parser and an Invoice Parser return structurally similar protos but with different field-level annotations.
Foxit’s PDF Structural Extraction API runs three coordinated layers on every document, regardless of document type. The OCR layer handles rasterized content across more than 200 languages. The layout recognition layer maps spatial relationships and table cell grids, resolving multi-column text blocks, overlapping text-image regions, stamped signatures on top of text fields, and engineering drawing annotations. The AI parsing layer then classifies content semantically, assigning each element a type from a fixed taxonomy of twelve labels.
The rendered pipeline. Every document takes the same path, so there is no processor to choose per document type.
Those twelve types appear in StructureInfo.json inside the returned ZIP archive, covering title, head, paragraph, table, image, headerFooter, form, hyperlink, footnote, sidebar, annotation, and formula. Each element carries its reading-order position and spatial coordinates, so the structure your downstream system receives reflects how a human would read the original document rather than raw storage order.
The practical delta shows up in RAG pipeline integration. Google Document AI’s Document proto gives you text and coordinates, but your pipeline needs a post-processing step to decide what each block means semantically. With StructureInfo.json, you filter to table elements, iterate rows, and pass the content directly to your embedding model, because the semantic classification happened upstream.
The output schema comparison below clarifies where each platform puts the interpretive work:
| Schema dimension | Google Document AI (Document proto) | Foxit (StructureInfo.json) |
|---|---|---|
| Structure model | Hierarchical, covering pages, blocks, paragraphs, and tokens | Flat list of semantically typed elements with spatial metadata |
| Semantic labels | Field-level labels tied to specific processor selection | 12 fixed element types, processor-independent |
| Table representation | Cell tokens within a table block | Dedicated table element type with cell grid coordinates |
| Reading order | Implicit (coordinate ordering required client-side) | Explicit, preserved in element sequence |
| Formula support | Limited | Dedicated formula element type |
Document coverage on both platforms is broad. Foxit processes scanned PDFs, image-based PDFs, multi-page contracts, invoices, and form-heavy documents. The layout layer handles edge cases that trip up simpler OCR tools, including stamped signatures overlapping text fields, mixed raster-vector pages, and footnote regions that appear spatially disconnected from their reference markers.
Data privacy and ecosystem independence
Regulated workloads ask two questions before any technical evaluation, starting with where the document goes and what audit evidence you can produce.
Foxit’s API compliance page documents SOC 2 Type II independent audit, HIPAA-aligned features with Business Associate Agreement (BAA) support, and GDPR-supporting features including redaction, anonymization, and secure metadata handling. Foxit’s AI service documentation states that input documents and results are held temporarily and deleted within 24 hours. If your organization requires a BAA, Foxit can provide one.
Google Document AI routes all processing through GCP infrastructure. Teams subject to data residency requirements need to select the appropriate GCP region, review Google’s data processing addendum, and confirm their cloud agreement covers the specific data types being processed. That review is standard for teams already operating within GCP, but it adds a compliance surface for teams that are not.
Foxit’s API is cloud-agnostic. Your team calls it from any existing stack, passing two credential headers, and extracts documents without spinning up a GCP project, provisioning a storage bucket, or accepting GCP billing terms. For teams evaluating outside GCP, that independence cuts both technical and commercial risk from the decision.
When to use Google Document AI and when to use Foxit
The right choice depends on your existing infrastructure and what your extraction output needs to do.
| Scenario | Best fit | Key reason |
|---|---|---|
| Teams already deep in GCP who want Vertex AI integration | Google Document AI | Native Vertex AI pipeline support and Google-managed processor versions reduce ops overhead |
| High-volume PDF processing outside the GCP ecosystem | Foxit | Cloud-agnostic with credit-based pricing, with no GCP billing or IAM dependency |
| Workloads with strict data residency or BAA requirements | Foxit | SOC 2 Type II audit, HIPAA-aligned features with BAA support, and GDPR-supporting features |
| Teams that need semantic element classification ready for downstream consumption | Foxit | StructureInfo.json delivers 12 pre-labeled element types without a client-side post-processing step |
The fourth scenario is worth walking through in detail. Your team has a contract review pipeline that needs to extract all tables and footnotes from multi-page PDFs and push them into a downstream system. With the Foxit API, the workflow runs like this:
- Upload the PDF to
/pdf-services/api/documents/uploadand receive adocumentId. - POST to
/pdf-services/api/documents/pdf-structural-extractwith thedocumentIdand receive ataskId. - Poll
/pdf-services/api/tasks/{taskId}every two seconds untilstatusequalsCOMPLETED. - Download the result ZIP from
/pdf-services/api/documents/{resultDocumentId}/downloadand parseStructureInfo.json.
Once you have StructureInfo.json, filtering to table and footnote elements is a single-pass list comprehension. The labeled elements arrive with spatial coordinates and reading order intact, so your downstream system receives structured, ordered data ready for indexing, embedding, or display, with no second model call, no coordinate sorting, and no block-level classification required.
Teams running active Vertex AI pipelines in GCP, with processors not among the seven being retired, have no technical reason to switch. For everyone else, the four API calls above give you a working extraction against your own documents in minutes.
Google Document AI FAQ
What is Google Document AI?
Google Document AI is a managed cloud service for document parsing and data extraction, built on GCP processors. You select a processor type (such as Invoice Parser, Form Parser, or Document OCR), send a document via the API, and receive a structured Document proto containing text, coordinates, and field-level annotations. All processing runs on Google Cloud Platform infrastructure.
How does the Foxit PDF Structural Extraction API differ from Google Document AI?
Foxit’s API runs three coordinated processing layers on every document (OCR, layout recognition, and AI parsing) regardless of document type, while Google Document AI uses a processor model where semantic classification depends on the specific processor you select. The output schemas also differ. Foxit returns StructureInfo.json inside a ZIP archive with twelve pre-labeled element types preserving reading order and spatial relationships, while Google Document AI returns a Document proto with hierarchical blocks, paragraphs, and tokens that require client-side semantic interpretation. Foxit is also cloud-agnostic and needs no GCP project, while Google Document AI is GCP-native.
What are Foxit’s data-handling and compliance commitments for the extraction API?
Foxit’s API compliance page documents SOC 2 Type II independent audit, HIPAA-aligned features with Business Associate Agreement (BAA) support, and GDPR-supporting features including redaction, anonymization, and secure metadata handling. Foxit’s AI service documentation states that input documents and results are held temporarily and deleted within 24 hours. Foxit does not claim HIPAA certification or GDPR certification, and the API compliance page does not state that documents are never stored, so teams with specific retention requirements should review the documentation directly and request a BAA where applicable.
What output format does the Foxit PDF Structural Extraction API return?
The API returns a ZIP archive containing StructureInfo.json. That file classifies every element in the document using one of twelve labeled types, including title, head, paragraph, table, image, headerFooter, form, hyperlink, footnote, sidebar, annotation, and formula. Each element includes spatial coordinates and reading-order position, so the structure your downstream system receives reflects how a human would read the original document. Table elements carry cell grid coordinates, making row-level data extraction straightforward without additional parsing.
Can I try the Foxit PDF Structural Extraction API without a paid plan?
Yes. Create a free developer account at app.developer-api.foxit.com/sign-up. No credit card is required. Once you register, your client_id and client_secret are available immediately in the Developer Portal, and you can run extractions against your own documents using the API Playground or the downloadable Postman collection.
How does Foxit handle scanned or image-based PDFs?
Foxit’s OCR layer processes rasterized content across more than 200 languages, so scanned PDFs and image-based documents (including TIFFs) are processable without pre-conversion. The layout recognition layer then maps spatial relationships and resolves edge cases such as multi-column text blocks, overlapping text-image regions, and footnote regions that appear spatially disconnected from their reference markers.
Which document types does Google Document AI support after the June 2026 deprecations?
After 30 June 2026, Google is retiring the Enterprise Document OCR, Expense, Custom classifier, Custom splitter, Invoice, Pay slip, and Bank statement processors. Remaining processors (including Form Parser and Document OCR for non-deprecated versions) continue to operate on their own rolling deprecation schedule, where each version is deprecated six months after a newer one ships. Teams relying on any of the seven retired processors need to migrate before that date.
Conclusion
Across all five axes, the practical difference between Google Document AI alternatives comes down to how much infrastructure you take on to reach a first result. Foxit requires two credential headers. Google’s setup adds a GCP project, service account, IAM configuration, and billing enablement before you process a single document. Foxit’s three-layer extraction model (OCR, layout recognition, AI parsing) delivers semantic classification across every document type without processor selection. StructureInfo.json‘s twelve labeled element types reduce client-side post-processing compared to Google’s Document proto. Both platforms handle the major document types, and Foxit’s 200-plus language OCR covers non-Latin scripts across all document categories. Foxit’s SOC 2 Type II audit, HIPAA-aligned BAA support, and GDPR-supporting features give regulated teams a documented compliance baseline, and the cloud-agnostic model means your team runs document extraction without taking on GCP billing, IAM governance, or ecosystem lock-in.
Create a free developer account at app.developer-api.foxit.com/sign-up, no credit card required, and run the four-step extraction against your own document today.
How to Turn PDFs into Structured Data with Foxit’s PDF Structural Extraction API

PDF data extraction with Foxit’s Structural Extraction API turns messy invoices and tables into typed JSON, complete with bounding regions and addressable cells. This tutorial walks through the four REST calls, upload, extract, poll, and download, and shows working Python code that builds a clean dictionary from an invoice’s line items. It also covers common mistakes like case-sensitive auth headers and stale document IDs.
Pull text out of a multi-column invoice and you get a flat string with column headers mixed into values, row boundaries gone, and field labels indistinguishable from the data they describe. Foxit’s PDF Structural Extraction API returns typed JSON instead, where every element carries a type, its text, a bounding region, and, for tables, an addressable grid of cells.
This tutorial walks the four REST calls that get you there, uploading a PDF, starting the analysis, polling the task, and downloading the result. By the end you’ll have working Python code that turns an invoice into a dictionary your pipeline can address by key.
Raw text vs. structured extraction
What separates raw text extraction from structured extraction is the shape of the output, not the accuracy of the characters.
Take a vendor invoice with a line-item table covering description, quantity, and unit price. Text extraction returns something like "1 API Integration Consulting 10 $ 150.00 $1,500.00". The content is all there, but the row and column relationships are gone, so your parsing code has to reconstruct structure the PDF already encoded, and it has to do that differently for every layout you encounter.
Structured extraction preserves what raw text discards. The pdf-structural-extract endpoint returns each element with a type, a content object holding the text and its font styling, and a region giving the page number and bounding polygon. Tables come back as a cell grid with explicit rowIndex and columnIndex values, so a cell’s position is data rather than something you infer from coordinates.
Prerequisites
- Python 3.8+ with pip and a virtual environment via venv.
- The requests library for the HTTP calls.
- curl if you want to try the endpoints before writing code.
- A code editor, VS Code with the Python extension is a good default, though PyCharm or Sublime Text work equally well.
- A Foxit Developer account, free with no credit card, created at app.developer-api.foxit.com/sign-up. Activate the Developer plan (500 credits per year) and copy the Client ID and Client Secret from the APIs Dashboard.
- A sample PDF, so you do not have to build one. This tutorial uses invoice_table_test.pdf, a one-page invoice with a five-column line-item table.
Scaffold the workspace in one shot:
mkdir foxit-extract && cd foxit-extract
python3 -m venv .venv && source .venv/bin/activate
pip install requests
curl -L -o invoice.pdf https://github.com/lucienchemaly/foxit-demo-templates/raw/main/invoice_table_test.pdf
export FOXIT_CLIENT_ID="your_client_id"
export FOXIT_CLIENT_SECRET="your_client_secret" Here is the invoice the rest of the tutorial extracts from.
The source document. The five-column table and the labeled fields above it are what the extraction turns into addressable JSON.
Authentication
PDF Services authenticates with two named request headers on every call, client_id and client_secret, both lowercase with an underscore. There is no OAuth exchange and no bearer token, and wrapping the credentials in an Authorization: Bearer header returns a 400 instead. The base host for every endpoint in this tutorial is https://na1.fusion.foxit.com/pdf-services.
Keep the values in environment variables rather than in the file, so nothing secret travels with your code.
The four-call extraction flow
Structural extraction is an asynchronous job, so it runs in four steps.
The four calls and what each one hands to the next. The id you download with comes from the finished task, not the upload.
- Upload the PDF and receive a
documentId. - Start the analysis against that id and receive a
taskId. - Poll the task until its
statusreachesCOMPLETED, which also returns aresultDocumentId. - Download the result, a ZIP archive holding the structured JSON.
Step 1: Upload the document
Send the PDF as multipart/form-data to the upload endpoint, using the form field name file.
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/upload" \
-H "client_id: $FOXIT_CLIENT_ID" \
-H "client_secret: $FOXIT_CLIENT_SECRET" \
-F "[email protected]" The -F flag is what makes curl send a multipart body, and the @ prefix tells it to read the file from disk rather than treat the value as a literal string. Uploads are capped at 100 MB, and an uploaded document is deleted after 24 hours, so treat the documentId as short-lived rather than a permanent handle.
A successful upload returns a single key:
{
"documentId": "6a6c9834a820c33d30d222e5"
} Step 2: Start the structural analysis
POST that id to the extraction endpoint with a JSON body.
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/pdf-structural-extract" \
-H "client_id: $FOXIT_CLIENT_ID" \
-H "client_secret: $FOXIT_CLIENT_SECRET" \
-H "Content-Type: application/json" \
-d '{"documentId": "6a6c9834a820c33d30d222e5"}' The call returns HTTP 202 with a taskId rather than the finished document, since analysis runs asynchronously. documentId is the only required field in the body, and a password-protected source PDF takes an optional password alongside it. Full request and response details live in the PDF Structural Extraction reference, which also carries the endpoint’s Trial designation, so pin the schema version you parse against rather than assuming it is stable.
{
"taskId": "6a6c9835d24a2429666f61b6"
} Step 3: Poll the task
Ask for the task by id until it finishes.
curl "https://na1.fusion.foxit.com/pdf-services/api/tasks/6a6c9835d24a2429666f61b6" \
-H "client_id: $FOXIT_CLIENT_ID" \
-H "client_secret: $FOXIT_CLIENT_SECRET" The response carries the state, a percentage, and, once the work is done, the id of the result document:
{
"taskId": "6a6c9835d24a2429666f61b6",
"status": "COMPLETED",
"progress": 100,
"resultDocumentId": "6a6c98375e2cab6bb50e740b"
} Task statuses are uppercase. The schema enum runs PENDING, IN_PROGRESS, COMPLETED, and FAILED, so a comparison against a lowercase "completed" never matches and your loop spins until it times out. Portal copy sometimes says “processing” in prose, but IN_PROGRESS is the value on the wire.
Step 4: Download the result
Fetch the finished artifact using the resultDocumentId from the poll, not the documentId from the upload. Confusing the two is the most common 4xx at this step.
curl -o extract.zip \
"https://na1.fusion.foxit.com/pdf-services/api/documents/6a6c98375e2cab6bb50e740b/download" \
-H "client_id: $FOXIT_CLIENT_ID" \
-H "client_secret: $FOXIT_CLIENT_SECRET" The response comes back as application/zip. Unzipping it gives you StructureInfo.json, the structured output, alongside a rendered PNG of each analyzed page (page_p0.pdf_0.png for a one-page file).
The whole flow in Python
Here is the complete script, reading credentials from the environment.
import os
import time
import zipfile
import json
import requests
BASE = "https://na1.fusion.foxit.com/pdf-services/api"
HEADERS = {
"client_id": os.environ["FOXIT_CLIENT_ID"],
"client_secret": os.environ["FOXIT_CLIENT_SECRET"],
}
def upload(path):
with open(path, "rb") as fh:
r = requests.post(f"{BASE}/documents/upload", headers=HEADERS, files={"file": fh})
r.raise_for_status()
return r.json()["documentId"]
def start_extract(document_id):
r = requests.post(
f"{BASE}/documents/pdf-structural-extract",
headers={**HEADERS, "Content-Type": "application/json"},
json={"documentId": document_id},
)
r.raise_for_status()
return r.json()["taskId"]
def wait_for_task(task_id, interval=3, timeout=180):
deadline = time.time() + timeout
while time.time() < deadline:
r = requests.get(f"{BASE}/tasks/{task_id}", headers=HEADERS)
r.raise_for_status()
body = r.json()
if body["status"] == "COMPLETED":
return body["resultDocumentId"]
if body["status"] == "FAILED":
raise RuntimeError(f"Extraction failed: {body.get('error')}")
time.sleep(interval)
raise TimeoutError(f"Task {task_id} unfinished after {timeout}s")
def download_zip(result_id, out="extract.zip"):
r = requests.get(f"{BASE}/documents/{result_id}/download", headers=HEADERS)
r.raise_for_status()
with open(out, "wb") as fh:
fh.write(r.content)
return out
document_id = upload("invoice.pdf")
result_id = wait_for_task(start_extract(document_id))
archive = download_zip(result_id)
with zipfile.ZipFile(archive) as z:
structure = json.loads(z.read("StructureInfo.json"))
analyze = structure["analyzeResult"]
print("schema", analyze["version"]["schema"], "pages", len(analyze["pages"])) In this code, you upload the invoice and keep the returned documentId, hand that id to the extraction endpoint to get a taskId, then poll the task on a fixed interval until it reports COMPLETED and yields a resultDocumentId. The download call writes the ZIP to disk, and rather than unpacking it to a folder you read StructureInfo.json straight out of the archive. The top-level key is analyzeResult, which is where the schema version, the page list, and the element array all live.
A three-second interval with a 180-second ceiling is comfortable for single-page documents. Back off rather than tightening the loop if you process long files, since polling every second only burns request budget without finishing the job sooner.
Reading the structured JSON
analyzeResult holds four things worth knowing about, an info block of document metadata, a version block, the pages array, and the elements array that carries the content.
{
"analyzeResult": {
"version": {
"schema": "1.0.7",
"software": "FoxitPDFAnalyzer",
"model": "idp-analysis"
},
"pages": [
{ "pageNumber": 1, "size": {}, "state": {} }
],
"elements": [
{
"type": "title",
"content": {
"text": "INVOICE",
"style": { "fontFamilyName": "Arial", "fontSize": 24.0 }
},
"region": {
"page": 1,
"boundingBox": [90, 71, 189, 71, 189, 99, 90, 99]
},
"score": 0.88,
"id": "title1"
}
]
}
} Each element follows the same shape. The type classifies it, and extracting this invoice returns title, head, paragraph, and table. The schema defines a wider set, adding image, headerFooter, form, hyperlink, footnote, sidebar, annotation, and formula, so branch on the types your documents actually produce rather than assuming only four exist. The text and its font styling sit under content, so you read content.text rather than a top-level text key. The region gives the one-based page plus a boundingBox, and that box is an eight-number polygon listing four corner pairs in order, not a four-number rectangle. Every element also carries a confidence score and a stable id such as title1 or paragraph2, and paragraphs additionally carry a paragraphOrder for reading sequence.
Tables are the interesting case. Rather than a headers array and a two-dimensional rows array, a table exposes a cell list under content.body:
{
"type": "table",
"content": {
"body": {
"rowCount": 4,
"columnCount": 5,
"cells": [
{
"paragraph": { "type": "paragraph", "content": { "text": "Description" } },
"rowSpan": 1,
"columnSpan": 1,
"rowIndex": 0,
"columnIndex": 1
}
]
}
}
} Each cell states its own rowIndex and columnIndex along with rowSpan and columnSpan, and its text lives at paragraph.content.text. That is more verbose than a plain grid, but it means merged cells stay describable and you never have to infer column membership from x-coordinates.
Turning the cell list into rows
Since the API hands you cells rather than rows, build the grid yourself once and work with it afterwards.
def table_to_grid(table):
body = table["content"]["body"]
grid = [["" for _ in range(body["columnCount"])] for _ in range(body["rowCount"])]
for cell in body["cells"]:
text = cell.get("paragraph", {}).get("content", {}).get("text", "")
grid[cell["rowIndex"]][cell["columnIndex"]] = text.replace("\r\n", " ")
return grid
tables = [e for e in analyze["elements"] if e["type"] == "table"]
header, *data_rows = table_to_grid(tables[0])
line_items = [dict(zip(header, row)) for row in data_rows]
text_blocks = {
e["id"]: e["content"].get("text", "")
for e in analyze["elements"]
if e["type"] in ("title", "head", "paragraph")
}
print(header)
for item in line_items:
print(item) The code above allocates an empty grid from rowCount and columnCount, then drops each cell’s text into its stated position, which sidesteps any assumption about cell ordering in the array. Cell text can contain literal \r\n where a label wraps inside its column, so the replace call flattens that to a space before it reaches your data layer. Splitting the first row off as the header lets you zip each remaining row into a dictionary keyed by column name, and the same comprehension pattern collects the title, heading, and paragraph text by element id.
Running it against the sample invoice prints the real extraction:
['#', 'Description', 'Qty', 'Unit Price', 'Line Total']
{'#': '1', 'Description': 'API Integration Consulting', 'Qty': '10', 'Unit Price': '$ 150.00', 'Line Total': '$1,500.00'}
{'#': '2', 'Description': 'Compliance Review', 'Qty': '5', 'Unit Price': '$ 200.00', 'Line Total': '$1,000.00'}
{'#': '', 'Description': '', 'Qty': '', 'Unit Price': 'Subtotal:', 'Line Total': '$2,500.00'} Two things in that output are worth designing around. The unit price arrives as the string "$ 150.00", so currency parsing is still your job, and the final row is a subtotal rather than a line item, which is a reminder that the analyzer reports table geometry rather than business meaning. Filter trailing rows on an empty # or Description before you treat them as products. If you want to inspect a full result without running the calls yourself, the StructureInfo_sample.json from this exact run is available to read.
Feeding the output into an agent or downstream workflow
Once the table is a list of dictionaries and the labeled text is keyed by id, the payload is already agent-ready.
agent_context = {
"document": {"schema": analyze["version"]["schema"], "pages": len(analyze["pages"])},
"text_blocks": text_blocks,
"line_items": line_items,
} Because every element also carries a region, you can layer spatial checks on top, such as confirming a table sits below a particular heading by comparing the y values in their bounding polygons before you trust the association. Foxit also publishes an MCP server for PDF Services, so the same operations are reachable from an agent that speaks the Model Context Protocol rather than raw HTTP.
Common mistakes
- PascalCase auth headers : the keys are lowercase
client_idandclient_secret.ClientIdandClientSecretdo not authenticate, and neither does anAuthorization: Bearerheader, which returns a 400. - Comparing status to a lowercase string : task statuses are uppercase, so test against
COMPLETEDandFAILED. - Downloading with the upload id : the download path takes the
resultDocumentIdfrom the completed task, not thedocumentIdfrom the upload. - Reading
elementsfrom the root : the array is nested underanalyzeResult, sostructure["analyzeResult"]["elements"]is the path. - Expecting
bboxor a rows array : positions arrive asregion.boundingBoxwith eight numbers, and tables arrive ascontent.body.cellswith index fields rather than a headers plus rows pair. - Treating the ZIP as the JSON : the download is an archive, and the structured output is the
StructureInfo.jsonentry inside it. - Reusing a stale
documentId: uploads are removed after 24 hours, and the cap on an upload is 100 MB. - Polling every second : that exhausts request budget without speeding anything up. A few seconds between checks is enough.
PDF data extraction FAQ
What element types does the structural extraction return?
Extracting this invoice produces title, head, paragraph, and table elements. Every element carries type, content, region, score, and id, with tables adding a cell grid under content.body.
How is this different from raw text extraction for tables?
Raw text collapses a table into one string and loses row and column boundaries. Structural extraction reports rowCount, columnCount, and a cell list where each cell states its own rowIndex and columnIndex, so position is data rather than inference.
Is the extraction synchronous?
No. The extract call returns HTTP 202 with a taskId, and you poll GET /pdf-services/api/tasks/{taskId} until the status reaches COMPLETED, which is when resultDocumentId appears.
What does the download actually contain?
A ZIP archive holding StructureInfo.json plus a rendered PNG per analyzed page. The JSON is the structured output and the PNG is useful for visual spot checks.
What schema version does the output use?
The sample run reports analyzeResult.version.schema of 1.0.7, produced by FoxitPDFAnalyzer with the idp-analysis model. Read the version from the payload rather than hardcoding it, since it can move.
Does a failed task tell me why?
The task object surfaces the failure state in status as FAILED, so branch on that and log the whole task body when it happens.
Can I extract several documents at once?
Each upload and each task is independent, so run them concurrently and keep one taskId per document. The upload cap is 100 MB per file.
Get started with Foxit’s PDF Structural Extraction API
The pattern is four calls. Upload the PDF for a documentId, start pdf-structural-extract for a taskId, poll until COMPLETED for a resultDocumentId, then download the ZIP and read StructureInfo.json. From there analyzeResult.elements gives you typed titles, headings, paragraphs, and a table cell grid you can turn into dictionaries in a few lines.
Create a free developer account (no credit card) at account.foxit.com/site/sign-up, grab your Client ID and Secret from the APIs Dashboard, and run the script above against invoice_table_test.pdf to see the structured output for yourself.
Building Agentic Document Workflows: How LLM Agents Use PDF APIs to Convert, Extract, and Sign at Scale

This guide walks through building agentic document workflows by exposing Foxit’s PDF and eSign APIs as callable MCP tools, so any compatible agent host can run the full document lifecycle in one automated pipeline.
Agentic document workflows go beyond retrieval, since they convert, transform, merge, and sign documents without human intervention. This guide shows how to expose Foxit’s production PDF API as callable MCP tools so any LLM agent can execute the full document lifecycle (OCR, extraction, generation, and legally binding signatures) in a single automated pipeline.
Most LLM-powered applications have solved the retrieval problem. The harder part of agentic document workflows is action, the moment your agent needs to convert a scanned invoice to searchable text, merge a dozen contract pages into a package, and route it for a legally binding signature without a human in the loop.
RAG gets text into a context window, which is useful for reading. Once you need to produce, transform, or sign a document, you’ve moved into document operations territory. A plain text API won’t close that delta, and bolting together a dozen bespoke REST wrappers every time you need a new pipeline quickly becomes the bottleneck.
The Model Context Protocol (MCP) gives agents a standard way to discover and call tools. Document workflows have been missing a tool surface that exposes real PDF operations as callable MCP tools, backed by a production-ready API. This guide walks through exactly how to build that.
What You Need Before You Start
Five prerequisites are required to follow this guide. You need a Foxit developer account, the open-source MCP server, an MCP-compatible host, three environment variables, and a Python workspace for the signing example.
A Foxit developer account. Sign up at account.foxit.com/site/sign-up (no credit card required for the free Developer plan). The Foxit Developer Portal issues your Client ID and Client Secret, gives you access to the API Playground, and tracks usage in real time.
The open-source MCP server. Clone github.com/foxitsoftware/foxit-pdf-api-mcp-server. The repo ships two active implementations, a Python build (using FastMCP, Python 3.11+, and the uv package manager) and a TypeScript build (Node.js 18+, pnpm). The original stdio-python variant is deprecated, so use the current Python or TypeScript implementation.
An MCP-compatible host. You need somewhere to run the agent. Claude Desktop, Cursor, or VS Code with GitHub Copilot all work, and any MCP-compliant custom agent framework will also connect to the server.
A Python workspace for the signing example. The eSign walkthrough later in this guide runs a short Python script, so you need Python 3.8+ and the requests library. That walkthrough also uses a separate set of eSign credentials, which you set up in its own section rather than here. Scaffold an isolated workspace in one shot:
mkdir agentic-docs && cd agentic-docs
python3 -m venv .venv && source .venv/bin/activate
pip install requests Three environment variables. Before launching your MCP host process, export these:
export FOXIT_CLOUD_API_HOST="https://na1.fusion.foxit.com/pdf-services"
export FOXIT_CLOUD_API_CLIENT_ID="your_client_id"
export FOXIT_CLOUD_API_CLIENT_SECRET="your_client_secret" Never hardcode credentials in config files. The MCP server reads these at startup and uses them to authenticate every request to the PDF Services API.
What “Agentic” Actually Means for Document Workflows
An agentic document workflow executes operations on documents (converting formats, applying OCR, merging pages, routing for signature) rather than simply retrieving text from them. The tool surface required is fundamentally different from a RAG setup.
Retrieval-augmented generation pulls text from a document and injects it into a prompt. An agentic document workflow does something to a document, whether it converts a format, applies OCR to make a scanned image searchable, merges pages from multiple sources, or routes the result for signature.
In a tool-use architecture, the LLM doesn’t call the API directly. It picks the right operation from a catalog of tools based on the task, calls it with structured inputs, processes the result, and decides whether to continue the chain or hand off to the next step. If you’ve worked with web-search or code-execution tools in LangChain or AutoGen, the pattern is identical. The model reasons about which tool to invoke, not about how the underlying HTTP request works.
A REST API is an HTTP surface. An MCP tool is a named, typed function with an input schema, an output contract, and a description the model uses to decide when and whether to call it. MCP standardizes that interface so any compliant host can discover the full tool catalog, call individual operations, and chain results without custom adapter code.
A well-designed MCP server eliminates the bespoke integration layer. Without one, every document-heavy agent pipeline requires someone to write and maintain that plumbing from scratch.
Architecture Overview: Two Modes for Agent-Driven PDF Processing
The Foxit PDF API MCP Server wraps Foxit’s cloud PDF Services API as 30+ callable MCP tools, covering every stage of a document lifecycle. Foxit PDF Editor is the first PDF editor in the industry to act as an MCP Host, connecting outward to external MCP Servers and acting on open documents.
Those two facts define two distinct architectural modes.
Mode 1: Programmatic pipeline. Your MCP host (Claude Desktop, Cursor, VS Code with GitHub Copilot, or a custom agent) registers the Foxit PDF API MCP Server. The agent calls PDF tools directly, the server translates those calls into Foxit PDF Services REST requests, and structured results return to the agent. The agent never writes REST plumbing. This is the right model for automated pipelines running without a human in the loop.
Mode 2: In-app orchestration. Foxit PDF Editor acts as the MCP Host. Its embedded AI Assistant connects to external MCP Servers (Jira, Salesforce, Gmail, Notion, GitHub, Google Workspace) and acts on the open document. You could extract fields from a contract PDF and open a Jira ticket without leaving the editor. This is the right model when a knowledge worker needs AI assistance during document review.
Mode 1 is what the rest of this guide builds. Its data flow runs like this:

The agent calls tools, the MCP server handles the REST layer against PDF Services, and a prepared document hands off to eSign at the end of the chain.
One detail to understand before you build is that each successful Foxit PDF Services API call consumes one credit from your plan. Failed requests (4xx or 5xx) do not consume credits. The Developer Dashboard shows real-time usage, so you can see exactly what a pipeline costs per document before scaling it up.
Setting Up the Foxit MCP Server
The server exposes tools across six categories (document lifecycle, creation, conversion, manipulation, security, and forms) plus OCR and document compare.
The full tool catalog breaks down as follows:
- Document lifecycle : upload, download, delete
- Creation : Word, Excel, PowerPoint, HTML, URL, plain text, and image to PDF
- Conversion : PDF to Word, Excel, PowerPoint, HTML, plain text, and image
- Manipulation : merge, split, extract pages, compress, flatten, linearize, watermark, and page operations
- Security : add and remove passwords, set permissions
- Forms : export and import form data as JSON
OCR and document compare are also in the catalog. Signing lives in the eSign API covered in Section 6.
Mode 1: Programmatic Pipeline Setup
Clone the repo and pick your implementation. For the Python version with VS Code and GitHub Copilot, create or update your .vscode/mcp.json with the following:
{
"servers": {
"foxit-pdf": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/foxit-pdf-api-mcp-server",
"run",
"foxit-pdf-api-mcp-server"
],
"env": {
"FOXIT_CLOUD_API_HOST": "${env:FOXIT_CLOUD_API_HOST}",
"FOXIT_CLOUD_API_CLIENT_ID": "${env:FOXIT_CLOUD_API_CLIENT_ID}",
"FOXIT_CLOUD_API_CLIENT_SECRET": "${env:FOXIT_CLOUD_API_CLIENT_SECRET}"
}
}
}
} VS Code launches the MCP server as a subprocess through uv, which runs the cloned Python build from its directory. The three ${env:...} references pull credentials from the shell environment instead of hardcoding them in the file, so replace /absolute/path/to/foxit-pdf-api-mcp-server with the path where you cloned the repo and restart VS Code to load the server. Claude Desktop uses the same shape under an mcpServers key and can run the published npm package directly with "command": "npx" and "args": ["-y", "@foxitsoftware/foxit-pdf-api-mcp-server"], so you do not have to clone anything for that host.
Mode 2: In-App Orchestration Setup
Open Foxit PDF Editor, click the AI Assistant tab in the Ribbon, and launch AI Chat to open the right-hand panel. In the bottom left of that panel, click MCP Tools, then click Add MCP Server to configure a new MCP service. Fill in the required fields and save. Once configured, the server appears in the MCP Tools list and its tools activate inside AI Chat.
In both modes, the server reads your Client ID and Client Secret from the environment variables set at startup. No additional gateway configuration is required.
Core Document Operations Agents Can Execute
With the server running, your agent has 30+ callable PDF operations available. Five categories do the heaviest lifting in production pipelines, namely conversion, OCR, structural extraction, merge/split, and document generation via DocGen.
Conversion. An agent receiving an uploaded DOCX triggers the Word-to-PDF creation tool before any downstream step. The output is a standards-compliant PDF that every subsequent operation (OCR, extraction, merge) can work with consistently, eliminating manual conversion and format ambiguity downstream.
OCR. When an agent ingests a scanned or image-only PDF, an OCR call should precede any extraction step. Calling the OCR tool makes the document searchable and text-extractable, which is required for invoice and contract pipelines where key fields sit inside scanned images. The agent calls OCR, waits for the result, and proceeds.
Structural extraction. After OCR, an agent extracts text, tables, and form-field data as structured JSON. Foxit’s structural extraction returns per-page content plus images, giving you a payload that routes cleanly into a BI tool or a second LLM step for analysis or classification. For the full response schema, refer to the Foxit MCP Server developer blog.
Merge and split. An agent assembling a contract package from multiple source documents calls the merge tool with an ordered list of PDFs. An agent pre-processing a large compliance document for parallel LLM analysis calls the split tool to divide it into per-section chunks. Both operations are synchronous and safe to retry on failure.
Document generation via DocGen. For dynamically generated contracts, invoices, or reports, the Foxit Document Generation API accepts a DOCX template with {{dynamic_tags}} and a JSON data payload from a CRM, database, or form response. It returns a finished PDF via POST /document-generation/api/GenerateDocumentBase64. A ready-to-use template lives in the Foxit demos repo if you want to test the call without authoring one. When you upload a DOCX template directly, the 4 MB post-base64 encoding cap applies, so slim templates down by stripping embedded fonts and large images before encoding. The agent supplies the JSON payload at runtime, so a single template can produce thousands of unique documents.
A representative end-to-end pipeline runs like this. An agent ingests a purchase order scan, calls OCR to make it searchable, extracts the structured fields as JSON, merges that data into a contract template via DocGen, and hands the finished PDF off to eSign. Every step is a tool call. The agent reasons about sequencing while the MCP server handles the REST layer.
Agent-Triggered Signing Workflows via the eSign API
Document signing uses a separate REST service, the Foxit eSign API, which has its own credentials and completes the pipeline in three calls. The agent exchanges its credentials for a token, creates a signing folder from the prepared PDF, and dispatches it to the signer.
The eSign API runs on its own host and issues its own Client ID and Client Secret from the eSign portal, separate from the PDF Services credentials the MCP server uses. Export the three eSign variables the script reads before running it:
export FOXIT_ESIGN_BASE_URL="https://na1.foxitesign.foxit.com"
export FOXIT_ESIGN_CLIENT_ID="your_esign_client_id"
export FOXIT_ESIGN_CLIENT_SECRET="your_esign_client_secret" The folder can only be sent if the signer has a signature field, and the simplest way to place one is with Foxit eSign text tags embedded in the document. Download the ready-to-sign sample, agent_agreement.pdf, into your workspace as agreement.pdf. It already carries the tag that maps a signature field to the first party, so the folder is sendable as is. Then run:
import base64
import os
import requests
BASE_URL = os.environ["FOXIT_ESIGN_BASE_URL"] # https://na1.foxitesign.foxit.com
CLIENT_ID = os.environ["FOXIT_ESIGN_CLIENT_ID"]
CLIENT_SECRET = os.environ["FOXIT_ESIGN_CLIENT_SECRET"]
def get_access_token():
resp = requests.post(
f"{BASE_URL}/api/oauth2/access_token",
data={
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
"grant_type": "client_credentials",
"scope": "read-write",
},
timeout=30,
)
resp.raise_for_status()
return resp.json()["access_token"]
def route_for_signature(pdf_path, signer):
token = get_access_token()
headers = {"Authorization": f"Bearer {token}", "Content-Type": "application/json"}
with open(pdf_path, "rb") as fh:
encoded = base64.b64encode(fh.read()).decode()
folder = requests.post(
f"{BASE_URL}/api/folders/createfolder",
headers=headers,
json={
"folderName": "Agent Service Agreement",
"inputType": "base64",
"base64FileString": [encoded],
"fileNames": ["agreement.pdf"],
"processTextTags": True,
"sendNow": False,
"parties": [
{
"permission": "FILL_FIELDS_AND_SIGN",
"firstName": signer["first_name"],
"lastName": signer["last_name"],
"emailId": signer["email"],
"sequence": 1,
}
],
},
timeout=60,
)
folder.raise_for_status()
folder_id = folder.json()["folder"]["folderId"]
sent = requests.post(
f"{BASE_URL}/api/folders/sendDraftFolder",
headers=headers,
json={"folderId": folder_id},
timeout=60,
)
sent.raise_for_status()
return folder_id
if __name__ == "__main__":
fid = route_for_signature(
"agreement.pdf",
{"first_name": "Jordan", "last_name": "Lee", "email": "[email protected]"},
)
print(f"Folder {fid} sent for signature") In this code, you read the eSign credentials from the environment and exchange them for a bearer token at the access_token endpoint, where the request is form-encoded rather than JSON (sending JSON returns a 415). You then base64-encode the local PDF and post it to createfolder with inputType set to base64 so the API reads the base64FileString array, with processTextTags set to True so the document’s text tags become a real signature field, and with sendNow set to False so the folder is created as a draft instead of emailing anyone immediately. The parties array names the signer with FILL_FIELDS_AND_SIGN permission, you read the new id at folder.folderId, and you pass it to sendDraftFolder, which dispatches the draft to the signer. To confirm the result, call GET /api/folders/viewActivityHistory?folderId={id}, which returns the activity log once the folder has been shared (a draft returns logs of a non-shared folder can not be viewed). Foxit uses “folder” throughout the eSign API, never “envelope.”
Compliance is built into the API layer. The Foxit eSign API supports eIDAS, the ESIGN Act, UETA, HIPAA, and GDPR, covering the requirements for agents operating in legal, healthcare, and finance contexts. No separate compliance infrastructure is required.
Production Considerations: Compliance, Cost, and Error Handling
The Foxit PDF Services and eSign APIs are SOC 2 Type II certified, with HIPAA BAA support, GDPR compliance, and CCPA coverage built in. Three additional concerns (credit consumption, idempotency, and async polling) determine whether a pipeline runs reliably at scale.
Credit consumption. The free Developer plan includes 500 credits per year (annual reset, no rollover). The Startup plan is $1,750/year for 3,500 credits. The Business plan is $4,500/year for 150,000 credits. Each successful API call consumes one credit; 4xx and 5xx responses do not. A pipeline processing 50 documents per day burns through those allocations quickly. Monitor real-time usage in the Developer Dashboard before moving to production and size your plan accordingly.
Idempotency. Merge, flatten, and convert calls are safe to retry on failure. Signing-folder creation is not, so gate createfolder behind a state check so the agent doesn’t create duplicate folders and dispatch duplicate signing requests to the same signers.
Async operations. Translation and batch conversion operations are asynchronous, so they return a job ID immediately, and your agent should poll the job status endpoint roughly every three seconds until the status is COMPLETED or FAILED before proceeding to the next step in the chain. Handling these concerns at build time separates a working pipeline from one that generates support tickets.
Common Mistakes
- Dropping the
/api/prefix on eSign calls : Every eSign path lives under/api/, as in/api/folders/createfolder. Omitting it returns a 404 against a docs-style path that does not exist. - Sending the token request as JSON : The
access_tokenendpoint is form-encoded. A JSON body returns415 Unsupported Media Type, so pass the credentials as form data. - Forgetting
inputType: base64: When you send a base64 PDF without it,createfolderrejects the request withfileUrls or base64FileString cannot be empty. URL mode usesfileUrlsandfileNamesinstead. - Sending a signer with no signature field : A
FILL_FIELDS_AND_SIGNparty needs a field. If the document has no text tag like${s:1:______}and you skipprocessTextTags,sendDraftFolderreturnsPlease assign a signature field. Use underscores in the tag placeholder, since an empty placeholder does not create a field. - Expecting a signing tool in the MCP server : The 30+ MCP tools cover PDF operations, not signatures. Signing is the eSign API, a separate service with separate credentials.
Agentic Document Workflows FAQ
What is an agentic document workflow?
An agentic document workflow is an automated pipeline in which an LLM agent executes operations on documents (conversion, OCR, extraction, merging, generation, and signing) without human intervention. Unlike retrieval-augmented generation, which only reads documents, an agentic workflow produces and transforms them using callable tools exposed through a protocol like MCP.
How does the Model Context Protocol (MCP) work with PDF APIs?
MCP defines a standard interface for exposing named, typed functions, called tools, that an LLM agent can discover and invoke. A PDF API MCP server wraps REST endpoints as MCP tools with input schemas and output contracts. The agent selects the right tool based on task context, calls it with structured parameters, and processes the result, without writing any HTTP request logic.
What PDF operations does the Foxit MCP server expose?
The Foxit PDF API MCP Server exposes 30+ tools covering document lifecycle (upload, download, delete), creation (Word, Excel, PowerPoint, HTML, image to PDF), conversion (PDF to multiple formats), manipulation (merge, split, compress, OCR, watermark), security (password management), and forms (JSON import/export).
Does every API call consume a credit even if it fails?
No. Only successful Foxit PDF Services API calls consume credits. Requests that return 4xx or 5xx status codes do not count against your plan. The Developer Dashboard provides real-time usage tracking so you can measure pipeline cost per document before scaling.
How do I trigger document signing from an agent without human intervention?
After preparing a document, your agent calls the Foxit eSign API directly. It authenticates with client_credentials, POSTs to /api/folders/createfolder with the base64 PDF and a parties entry for the signer, then POSTs to /api/folders/sendDraftFolder with the returned folderId. The signing email workflow triggers automatically, and the agent can poll /api/folders/viewActivityHistory for the audit trail.
What compliance standards does the Foxit eSign API meet?
The Foxit eSign API supports eIDAS, the U.S. ESIGN Act, UETA, HIPAA, GDPR, and CCPA. The PDF Services API is SOC 2 Type II certified with HIPAA BAA support. No separate compliance wrappers are required for agents operating in legal, healthcare, or financial contexts.
What is the difference between Mode 1 and Mode 2 in the Foxit MCP architecture?
Mode 1 is a programmatic pipeline where an external LLM agent (Claude Desktop, Cursor, VS Code with GitHub Copilot, or a custom framework) registers the Foxit MCP Server and calls PDF tools automatically. Mode 2 is in-app orchestration where Foxit PDF Editor acts as the MCP Host, connecting to external services like Jira or Salesforce while a knowledge worker reviews the open document.
Why should signing-folder creation be gated behind a state check?
The /api/folders/createfolder endpoint is not idempotent. If an agent retries on failure without a state check, it will create duplicate folders and send duplicate signing requests to the same signers. Merge, flatten, and convert operations are safe to retry; folder creation requires the agent to verify no existing folder was created before calling the endpoint again.
Start Building: Free Developer Access in Minutes
The full pipeline in this guide is available to test today. Activate a free Foxit Developer plan to get your Client ID, Client Secret, access to the API Playground, and 500 credits for real requests. No credit card required.
Clone the open-source MCP server, export the three environment variables, and register the server in Claude Desktop, Cursor, or VS Code with GitHub Copilot. At that point, 30+ PDF tools are callable from your agent with no local SDK to install and no REST plumbing to write.
Once a document is prepared, extend the pipeline into legally binding signatures with the eSign API. The architecture in this guide covers the full document lifecycle from conversion and OCR through extraction, generation, merging, and signing, all in a single agentic document workflow.
Create your free developer account to clone the open-source MCP server and make 30+ PDF tools callable from your agent in minutes, no credit card required.
Programmatic PDF Editing with Foxit PDF Services API: Pages, Merges, Splits, and Flattening at Scale

Manually editing PDFs doesn’t scale when you’re processing hundreds of documents a day. This guide uses working Python and cURL examples to walk through the Foxit PDF Editor API’s core operations, covering page manipulation, merging, splitting, and flattening.
If you’re building document workflows at scale, you already know that manual PDF editing doesn’t cut it. Whether you’re generating contracts, processing invoices, or packaging reports, you need an API that handles page manipulation, merging, splitting, and flattening without breaking under load. Foxit PDF Services API gives you exactly that.
This guide covers the operations you’ll use most: adding and removing pages, merging documents, splitting by page range, and flattening annotations and form fields. Code examples are in Python, but the API is REST-based, so the patterns translate to any stack.
What Is the Foxit PDF Services API?
Foxit PDF Services API is a cloud-based REST API for programmatic PDF manipulation. You send documents and parameters; it returns processed PDFs. No local dependencies, no rendering engine to maintain.
The API handles:
Page insertion, deletion, and reordering
Document merging (combining multiple PDFs into one)
Document splitting (breaking one PDF into multiple outputs)
Flattening (converting annotations, form fields, and overlays into static page content)
Authentication uses API keys. Every request requires client_id and client_secret as separate HTTP request headers.
Why Backend PDF Editing Is an Infrastructure Decision
Teams processing hundreds of PDFs a day have an architecture problem, not a tooling problem.
When your CRM spits out contracts, your ERP generates invoices, and your intake forms produce patient records, the question is how you process these at scale without standing up a server farm or babysitting a library version matrix. At 500+ documents per day, manual desktop tooling is off the table. Per-file scripts that depend on a locally installed library become a liability the moment the library version shifts or a new language target appears.
Two architectural approaches dominate this space: SDK-based libraries and cloud REST APIs. SDK libraries require local installation, version pinning, language-specific bindings, and ongoing maintenance every time a dependency shifts. A cloud REST API requires none of that. Any language that can send an HTTP request can call it, with no package to install, no runtime to configure, and no compatibility matrix to manage.
This guide covers four operations in full: page manipulation (move, rotate, delete, add), merging multiple PDFs into one output, splitting a large document by page count, and flattening annotations into permanently static content. All of it runs via REST calls against the Foxit PDF Services API, authenticated with two request headers and structured around a single four-step loop that applies to every endpoint in the suite.
Prerequisites
Get your environment in order before any code runs. Each tool links to its canonical install page.
venv for project isolation
Node.js 18+ and npm
axiosfor Node examplesA code editor such as VS Code with the Python extension; alternatives include PyCharm, WebStorm, and Sublime Text
Postman (optional, for API exploration alongside the Foxit API Playground)
A Foxit developer account, available at account.foxit.com/site/sign-up with no credit card required and free credits included
Set up your workspace:
mkdir foxit-pdf-tutorial && cd foxit-pdf-tutorial
python3 -m venv .venv
source .venv/bin/activate
pip install requests
Authentication and the Upload → Task → Poll → Download Loop
Authentication
Foxit PDF Services API authenticates via headers. Every request must include client_id and client_secret as separate HTTP request headers, exactly as named: lowercase, underscored. They’re passed individually, as plain strings. Concatenating them, Base64-encoding them, or prefixing them with Bearer are all auth mistakes that look like they should work and don’t.
The base host for the North America environment is https://na1.fusion.foxit.com. Set it once as an environment variable and reuse it:
export BASE_URL="https://na1.fusion.foxit.com"
export CLIENT_ID="your_client_id"
export CLIENT_SECRET="your_client_secret" Every cURL request in this guide uses that pattern:
--header "client_id: $CLIENT_ID"
--header "client_secret: $CLIENT_SECRET"
In Python, read credentials from os.environ and never hardcode them:
import os
import time
import requests
BASE_URL = os.environ["BASE_URL"]
CLIENT_ID = os.environ["CLIENT_ID"]
CLIENT_SECRET = os.environ["CLIENT_SECRET"]
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
} In Node.js:
const axios = require("axios");
const BASE_URL = process.env.BASE_URL;
const headers = {
client_id: process.env.CLIENT_ID,
client_secret: process.env.CLIENT_SECRET,
}; The Four-Step Loop
Every operation in this API follows the same four steps, regardless of endpoint.
POST /pdf-services/api/documents/uploadwith the file asmultipart/form-data. The response returns an uploaddocumentId.POSTto the relevant endpoint (manipulate, combine, split, or flatten) with the uploaddocumentIdin the request body. The response is HTTP 202 with ataskId.GET /pdf-services/api/tasks/{task-id}until thestatusfield readsCOMPLETED. The completed task response also carries aresultDocumentIdand aprogresspercentage. Task statuses are always uppercase:PENDING,PROCESSING,COMPLETED,FAILED.GET /pdf-services/api/documents/{resultDocumentId}/downloadto retrieve your result. The uploaddocumentIdfrom step one and theresultDocumentIdfrom step three are different identifiers. You download using theresultDocumentId. Swapping these is the most common 4xx error.
Use a-midsummer-nights-dream.pdf to follow along with a known-good file:
curl --location "$BASE_URL/pdf-services/api/documents/upload" \
--header "client_id: $CLIENT_ID" \
--header "client_secret: $CLIENT_SECRET" \
--form 'file=@"a-midsummer-nights-dream.pdf"' Response:
{
"documentId": "abc123-upload-id"
}
A reusable Python polling function with exponential backoff:
def poll_task(task_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/tasks/{task_id}"
delay = 2
while True:
response = requests.get(url, headers=headers)
response.raise_for_status()
data = response.json()
status = data["status"]
if status == "COMPLETED":
return data["resultDocumentId"]
elif status == "FAILED":
raise RuntimeError(f"Task {task_id} failed: {data}")
time.sleep(delay)
delay = min(delay * 2, 30) Branch on "COMPLETED" and "FAILED" in uppercase. Lowercase strings will silently never match.
The async model is what makes this scalable. A 50-page merge or a batch of 200 splits doesn’t block your thread while it runs. You dispatch the operation call, collect the taskId, and poll on a sensible backoff schedule. At 500+ PDFs per day, that non-blocking pattern keeps your worker pool from saturating.
Page Manipulation: Move, Rotate, Delete, and Add Pages
All page operations go through a single endpoint: POST /pdf-services/api/documents/modify/pdf-manipulate.
{
"documentId": "<upload-document-id>",
"password": "optional",
"config": {
"operations": [{ "type": "OPERATION_TYPE" }]
}
} The config.operations array runs in order. Page indexing is 1-based and adjusts after each step. If you delete page 3 and then reference page 4 in the next operation, that reference points to what was originally page 5 before the delete. Keep that in mind when chaining operations in a single request.
MOVE_PAGES
MOVE_PAGES reorders pages within the document:
{
"type": "MOVE_PAGES",
"pages": [5, 6, 7],
"targetPosition": 1
} targetPosition is 1-based and must not exceed the document’s total page count.
ROTATE_PAGES
ROTATE_PAGES changes page orientation:
{
"type": "ROTATE_PAGES",
"pages": [1, 2, 3],
"rotation": "ROTATE_CLOCKWISE_90"
} Valid rotation values are ROTATE_0, ROTATE_CLOCKWISE_90, ROTATE_180, and ROTATE_COUNTERCLOCKWISE_90.
DELETE_PAGES
DELETE_PAGES removes specific pages:
{
"type": "DELETE_PAGES",
"pages": [8, 9]
} ADD_PAGES
ADD_PAGES appends blank pages to the end of the document:
{
"type": "ADD_PAGES",
"pageCount": 2
}
ADD_PAGES has no insert-at-position parameter; blank pages always go to the end. There’s also no REPLACE_PAGES operation in the current API. Page replacement requires a DELETE_PAGES call on the target pages followed by a separate merge step to bring in the replacement content.
A document intake pipeline for scanned medical records often receives pages in landscape orientation when the archive standard is portrait. The Python call below normalizes pages 1 through 3 of the uploaded file:
def rotate_pages(document_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/modify/pdf-manipulate"
payload = {
"documentId": document_id,
"config": {
"operations": [
{
"type": "ROTATE_PAGES",
"pages": [1, 2, 3],
"rotation": "ROTATE_CLOCKWISE_90"
}
]
}
}
response = requests.post(url, json=payload, headers=headers)
response.raise_for_status()
return response.json()["taskId"] A legal document workflow that assembles exhibits and needs to move pages 5, 6, and 7 to the front before archiving uses this request body:
{
"documentId": "<upload-document-id>",
"config": {
"operations": [
{
"type": "MOVE_PAGES",
"pages": [5, 6, 7],
"targetPosition": 1
}
]
}
}
All four operation types return HTTP 202 with a taskId. Pass that ID to poll_task, wait for COMPLETED, then download the resultDocumentId.
Merging and Splitting PDFs
Merging Multiple PDFs with pdf-combine
The merge endpoint is POST /pdf-services/api/documents/enhance/pdf-combine.
The request body takes a required documentInfos array. Each entry is a source document object:
{
"documentInfos": [
{ "documentId": "<doc-id-1>" },
{ "documentId": "<doc-id-2>" },
{ "documentId": "<doc-id-3>" }
],
"config": {
"addBookmark": true,
"continueMergeOnError": false,
"retainPageNumbers": false
}
} The config object supports three keys:
addBookmarkgenerates one bookmark per source file in the merged output, useful when the reader needs to navigate by source.retainPageNumberspreserves original page-number labels from each source.continueMergeOnErroris the one that matters most in production. Set it tofalsefor any job where a single bad source must fail the entire batch. Set it totrueonly for best-effort pipelines where partial output is acceptable. Always set it explicitly rather than relying on defaults.
The following code assembles a monthly client report from three sources. Download input.pdf, second.pdf, and input_for_compare.pdf to run this end-to-end:
def upload_file(path: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/upload"
with open(path, "rb") as f:
response = requests.post(url, headers=headers, files={"file": f})
response.raise_for_status()
return response.json()["documentId"]
def merge_pdfs(document_ids: list) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/enhance/pdf-combine"
payload = {
"documentInfos": [{"documentId": doc_id} for doc_id in document_ids],
"config": {
"addBookmark": True,
"continueMergeOnError": False,
"retainPageNumbers": False
}
}
response = requests.post(url, json=payload, headers=headers)
response.raise_for_status()
return response.json()["taskId"]
# Upload all three source files, collect their documentIds
doc_ids = [
upload_file("input.pdf"),
upload_file("second.pdf"),
upload_file("input_for_compare.pdf")
]
# Merge and download
merge_task_id = merge_pdfs(doc_ids) # POST combine → taskId
merge_result_id = poll_task(merge_task_id) # poll → resultDocumentId
Once poll_task returns a resultDocumentId, pass it to your download function.
Splitting a PDF by Page Count with pdf-split
The split endpoint is POST /pdf-services/api/documents/modify/pdf-split. The pdf- prefix is required and matches the naming convention used by pdf-combine, pdf-manipulate, and pdf-flatten.
The request body is minimal:
{
"documentId": "<upload-document-id>",
"pageCount": 10
} pageCount specifies how many pages go into each output file. The last file gets whatever pages remain. The current API exposes only this page-count split mode. File-size-based and bookmark-based splitting are not in the published reference.
curl --location "$BASE_URL/pdf-services/api/documents/modify/pdf-split" \
--header "client_id: $CLIENT_ID" \
--header "client_secret: $CLIENT_SECRET" \
--header 'Content-Type: application/json' \
--data '{
"documentId": "<upload-document-id>",
"pageCount": 10
}' The response returns a taskId. When polling reaches COMPLETED, the result contains multiple output files, one per split chunk.
Flattening Annotations: Making PDF Changes Permanent
Flattening merges annotations, form fields, and layers into the page content itself. The output is a single static PDF where no markup can be toggled, filled, or removed. This is a one-way transformation with no undo.
The endpoint is POST /pdf-services/api/documents/modify/pdf-flatten.
The request body is intentionally minimal:
{
"documentId": "<upload-document-id>"
} No flags for selective annotation or form-field handling are exposed in the current API reference. The operation flattens both annotations and form fields together in a single pass.
When to Flatten a PDF
Flattening is required in three situations:
Archiving a signed form. Without flattening, a technically capable recipient can still manipulate form fields in most PDF viewers, even after signature. Flattening removes that possibility entirely.
Print production. Annotation layers render differently across print drivers and produce visible artifacts in the final output. Flattening eliminates the variable.
Compliance workflows under HIPAA or legal hold. Documents at rest must be immutable. A live form field fails that requirement.
curl --location "$BASE_URL/pdf-services/api/documents/modify/pdf-flatten" \
--header "client_id: $CLIENT_ID" \
--header "client_secret: $CLIENT_SECRET" \
--header 'Content-Type: application/json' \
--data '{ "documentId": "<upload-document-id>" }' Flattening does not encrypt the document or apply password protection. If your workflow requires both, chain a pdf-protect call after flattening using POST /pdf-services/api/documents/security/pdf-protect, passing the resultDocumentId from the completed flatten task as the documentId input for the protect call and the password under config.userPassword. The result of one operation becomes the input of the next.
{
"documentId": "<flatten-result-document-id>",
"config": {
"userPassword": "<password>"
}
} In this body, you point documentId at the flatten task’s resultDocumentId and set the open password under config.userPassword. The password lives inside the config object, not at the top level. A top-level password key is accepted by the request but the task then fails with a parameter error, so nest it under config.
Building a Multi-Step PDF Editing Pipeline for High Volume
Chaining Operations with resultDocumentId
Every completed task returns a resultDocumentId. That ID becomes the documentId for the next operation. The full chain: upload, then operation A, then poll until A reaches COMPLETED, then operation B using A’s resultDocumentId, then poll until B reaches COMPLETED, then download B’s resultDocumentId.
A complete flatten-then-protect pipeline, with inline comments showing which ID type is active at each step:
import os
import time
import requests
BASE_URL = os.environ["BASE_URL"]
CLIENT_ID = os.environ["CLIENT_ID"]
CLIENT_SECRET = os.environ["CLIENT_SECRET"]
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
}
def upload_file(path: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/upload"
with open(path, "rb") as f:
response = requests.post(url, headers=headers, files={"file": f})
response.raise_for_status()
return response.json()["documentId"]
def poll_task(task_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/tasks/{task_id}"
delay = 2
while True:
response = requests.get(url, headers=headers)
response.raise_for_status()
data = response.json()
status = data["status"]
if status == "COMPLETED":
return data["resultDocumentId"]
elif status == "FAILED":
raise RuntimeError(f"Task failed: {data}")
time.sleep(delay)
delay = min(delay * 2, 30)
def flatten(document_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/modify/pdf-flatten"
response = requests.post(url, json={"documentId": document_id}, headers=headers)
response.raise_for_status()
return response.json()["taskId"]
def protect(document_id: str, password: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/security/pdf-protect"
payload = {"documentId": document_id, "config": {"userPassword": password}}
response = requests.post(url, json=payload, headers=headers)
response.raise_for_status()
return response.json()["taskId"]
def download(document_id: str, output_path: str):
url = f"{BASE_URL}/pdf-services/api/documents/{document_id}/download"
response = requests.get(url, headers=headers)
response.raise_for_status()
with open(output_path, "wb") as f:
f.write(response.content)
# Pipeline: upload → flatten → protect → download
upload_doc_id = upload_file("signed-form.pdf") # documentId
flatten_task_id = flatten(upload_doc_id) # taskId
flatten_result_id = poll_task(flatten_task_id) # resultDocumentId
protect_task_id = protect(flatten_result_id, os.environ["PDF_PASSWORD"]) # taskId
protect_result_id = poll_task(protect_task_id) # resultDocumentId
download(protect_result_id, "final-protected.pdf") Three distinct IDs flow through this pipeline: the upload documentId, the taskId returned by each operation call, and the resultDocumentId returned by each completed task. Confusing any two of these produces a 4xx.
Batch Polling for High-Volume Jobs
For a single document, sequential polling works fine. For a batch of thousands, polling each task one at a time becomes the bottleneck. The pattern that scales is to dispatch first, then consolidate.
Upload all source files in parallel and collect their documentIds. Dispatch all operation calls in parallel and collect their taskIds. Then run a single consolidated polling loop that iterates over all taskIds, checks each status, and branches on COMPLETED or FAILED per task. Your thread never blocks waiting for one document while others are already done, and it processes results as they arrive.
The multi.py sample in the foxitsoftware/developerapidemos repository demonstrates this exact dispatch-then-consolidate shape, including the resultDocumentId handoff across operations.
Common Mistakes and Troubleshooting
These are the integration errors that come up most often.
Polling too aggressively. Start at a two-second interval and double it up to a ceiling of around 30 seconds. Tight-looping the /tasks/{task-id} endpoint burns rate-limit budget and produces 429s that slow the whole pipeline down.
Treating FAILED as transient. A FAILED status means the job encountered a specific error. Read the task response body, surface the reason, and branch on it. Retrying indefinitely on FAILED produces the same failure indefinitely.
Downloading before COMPLETED. Some integrations skip the status check and call the download endpoint immediately after the operation response. The result is a partial or empty file. Always confirm status == "COMPLETED" before hitting the download endpoint.
Mixing up the three IDs. All three appear in the same code block so they can’t be confused:
upload_doc_id = upload_file("input.pdf") # documentId → input to operation endpoints
task_id = flatten(upload_doc_id) # taskId → input to the polling endpoint
result_doc_id = poll_task(task_id) # resultDocumentId → input to the download endpoint
download(result_doc_id, "output.pdf") The download endpoint takes the resultDocumentId, not the upload documentId or the taskId.
Task status casing. The polling endpoint returns PENDING, PROCESSING, COMPLETED, and FAILED in uppercase. Branching on lowercase "completed" will silently never match and your loop will run forever.
continueMergeOnError semantics. Setting this to true lets the merge skip a failing source and continue. Setting it to false aborts the entire batch if any source fails. In production, always set this explicitly.
Retrying 4xx responses. Retry on 5xx errors and timeouts. A 400 or 401 will return the same 4xx until you fix the request, so don’t retry them.
Auth header format. client_id and client_secret are separate headers, lowercase with underscores, passed individually. They’re not concatenated, Base64-encoded, or Bearer-prefixed.
pdf-flatten does not encrypt. Flattening makes a document’s content static but doesn’t restrict access to the file. If your compliance workflow requires both, chain a pdf-protect call after the flatten step using the flatten task’s resultDocumentId as the next documentId input.
PDF Editor API FAQ
What is the Foxit PDF Services API used for?
The Foxit PDF Services API is a cloud-based REST API for programmatic PDF manipulation at scale, covering page manipulation, document merging, splitting, and annotation flattening. Any language that can send HTTP requests can use it without installing local dependencies.
How do you merge multiple PDFs using the Foxit API?
Upload each source file to get a documentId, then POST all documentIds in a documentInfos array to POST /pdf-services/api/documents/enhance/pdf-combine. Poll the returned taskId until status is COMPLETED, then download the resultDocumentId.
What is PDF flattening and when should you use it?
PDF flattening converts interactive elements (form fields, annotations, and overlays) into static page content that cannot be edited. Use it when archiving signed forms, preparing documents for print production, or meeting compliance requirements such as HIPAA that mandate immutable records at rest.
What is the difference between documentId and resultDocumentId in the Foxit API?
documentId is returned after uploading a file and is used as input to operation endpoints. resultDocumentId is returned when a task reaches COMPLETED status and is used to download the processed file. They are different identifiers; using one where the other is expected produces a 4xx error.
Does the Foxit PDF Services API support splitting PDFs by custom page ranges?
The current pdf-split endpoint splits by a fixed pageCount value, producing equal-sized chunks with the last file containing remaining pages. File-size-based and bookmark-based splitting are not available in the published API reference.
How should you handle errors in the Foxit PDF Services API?
Retry on 5xx errors and timeouts using exponential backoff. Do not retry 400 or 401 responses, since those indicate a problem with the request itself. Read the response body on FAILED task status and branch on the specific error rather than retrying blindly.
Getting Started with Foxit PDF Services API
Create a free developer account. No credit card required, and free credits are included.
Once you have your credentials, open the Foxit API Playground. Start with pdf-combine. Download input.pdf and second.pdf, upload both to get two documentIds, POST the merge request with the JSON body from the Merging section, and poll for the result. The full cycle, authentication through download, takes under ten minutes on a first attempt.
If you’re evaluating this for production volume, check the current plan options and credit pools. Plans range from the free developer tier up through Startup and Business options with larger shared-credit allocations.
For the authoritative reference on every endpoint and parameter covered in this guide, the Foxit PDF Services API documentation is your starting point. The GitHub samples repository has complete Python and Node.js examples you can clone and run immediately, including the multi-step pipeline pattern this guide covers.
DOCX to PDF via the Foxit PDF Services API: Python and cURL Walkthrough

This walkthrough covers the full DOCX-to-PDF flow on the Foxit PDF Services API with runnable Python and cURL for each call.
Automated document pipelines demand conversion tooling that accepts a file, queues a job, and returns clear status at every step. The Foxit PDF Services API gives you exactly that: a four-endpoint async flow covering upload, convert, poll, and download. Each step returns a typed payload, the task model exposes four explicit states with a numeric progress field, and error codes map cleanly to distinct recovery paths.
This tutorial walks through every step of that flow in Python 3 with the requests library, plus cURL equivalents for each call. You’ll have a runnable convert.py script you can drop into a pipeline today.
Prerequisites
Before you run a single line of this tutorial, get the following in place. Each item links to its canonical install or setup guide.
- Python 3.8 or newer — verify with
python3 --version. The script uses only standard library modules plus one external package, so any modern 3.x will do. - pip — bundled with Python 3.4+. Verify with
python3 -m pip --version. - A virtual environment — isolates project dependencies so they don’t collide with system Python or other projects. See the venv tutorial for platform-specific activation commands.
- The
requestslibrary — the only third-party dependency in this walkthrough. Installed inside the venv below. - A code editor — Visual Studio Code with the Python extension is a solid default, but PyCharm, Sublime Text, or any editor you like will work.
- cURL — pre-installed on macOS and most Linux distros. Windows users can install from the official site or use WSL.
- A Foxit Developer account — register for free (no credit card required). The Foxit Developer Portal provisions a default application with your
CLIENT_IDandCLIENT_SECRETimmediately after signup.
Set up the project workspace:
mkdir foxit-docx-to-pdf && cd foxit-docx-to-pdf
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install requests Export your credentials as environment variables so the script never sees them as hardcoded strings:
export CLIENT_ID=your_client_id_here
export CLIENT_SECRET=your_client_secret_here import os
CLIENT_ID = os.environ.get("CLIENT_ID")
CLIENT_SECRET = os.environ.get("CLIENT_SECRET")
BASE_URL = "https://na1.fusion.foxit.com" All four API calls go to https://na1.fusion.foxit.com. The developer portal also offers a live sandbox and pre-built Postman collections if you want to verify calls in a GUI before scripting.
For a sample DOCX to work with right away, download input.docx directly from the foxitsoftware/developerapidemos GitHub repository and save it to your working directory.
How the Auth Model Works
The Foxit PDF Services API authenticates through named request headers. Pass client_id and client_secret directly on every call:
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
} json_headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
"Content-Type": "application/json",
} The API expects the raw key/secret pair in those named headers. Wrapping credentials in an Authorization: Bearer header instead returns 400, since the required client_id and client_secret headers are missing.
Step 1 and Step 2: Upload the DOCX and Initiate Conversion
Step 1: Upload the DOCX File
POST /pdf-services/api/documents/upload accepts the file as multipart/form-data and returns a documentId that every subsequent call needs.
import requests
def upload_doc(file_path: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/upload"
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
}
with open(file_path, "rb") as f:
files = {"file": (os.path.basename(file_path), f)}
response = requests.post(url, headers=headers, files=files)
response.raise_for_status()
return response.json()["documentId"] cURL equivalent:
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/upload" \
-H "client_id: $CLIENT_ID" \
-H "client_secret: $CLIENT_SECRET" \
-F "[email protected]" Uploaded files carry a 100 MB cap and are automatically deleted after 24 hours. A documentId scopes to the current upload session and expires with the source file, so treat it as ephemeral.
Step 2: Initiate the PDF Conversion
POST /pdf-services/api/documents/create/pdf-from-word accepts a JSON body with the documentId and returns a taskId. The API handles 10 to 10,000+ conversions per day across production pipelines, queuing jobs asynchronously to avoid blocking the connection until the PDF is ready.
import json
def convert_to_pdf(document_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/create/pdf-from-word"
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
"Content-Type": "application/json",
}
payload = {"documentId": document_id}
response = requests.post(url, headers=headers, data=json.dumps(payload))
response.raise_for_status()
return response.json()["taskId"] cURL equivalent:
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/create/pdf-from-word" \
-H "client_id: $CLIENT_ID" \
-H "client_secret: $CLIENT_SECRET" \
-H "Content-Type: application/json" \
-d '{"documentId": "<your_document_id>"}' The endpoint returns 202 Accepted, confirming the job is queued. It also accepts .doc, .rtf, .dot, .dotx, .docm, .dotm, and .wpd files through the same documentId input, so legacy Word formats work through the same pipeline.
Step 3: Polling the Task Status
GET /pdf-services/api/tasks/{task-id} returns four fields you need to act on in your polling loop:
status: one ofPENDING,IN_PROGRESS,COMPLETED, orFAILEDprogress: int32, 0 to 100resultDocumentId: populated when status reachesCOMPLETEDerror: populated when status reachesFAILED
The task state machine advances in one direction: PENDING to IN_PROGRESS, then to either COMPLETED or FAILED.

import time
def poll_task(task_id: str, max_attempts: int = 30) -> str:
url = f"{BASE_URL}/pdf-services/api/tasks/{task_id}"
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
}
for attempt in range(max_attempts):
response = requests.get(url, headers=headers)
response.raise_for_status()
data = response.json()
status = data.get("status")
progress = data.get("progress", 0)
print(f"Attempt {attempt + 1}: status={status}, progress={progress}%")
if status == "COMPLETED":
return data["resultDocumentId"]
if status == "FAILED":
raise RuntimeError(f"Conversion failed: {data.get('error')}")
time.sleep(2)
raise TimeoutError(f"Task {task_id} did not complete in {max_attempts} attempts") Two-second polling intervals work across a wide range of document sizes, and polling more aggressively only consumes rate limit budget without affecting conversion time.
Step 4: Downloading the Converted PDF
GET /pdf-services/api/documents/{documentId}/download fetches the finished PDF. The path parameter in the API reference reads {documentId}, but the value you pass here is the resultDocumentId from the completed poll response. The server assigns that ID to the generated PDF output at conversion time, making it the correct identifier to use at this step.
Stream the response to disk with stream=True and iter_content(chunk_size=8192). Buffering a large PDF fully into memory before writing it causes problems on high-volume pipelines.
def download_result(result_document_id: str, output_path: str) -> None:
url = f"{BASE_URL}/pdf-services/api/documents/{result_document_id}/download"
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
}
with requests.get(url, headers=headers, stream=True) as response:
response.raise_for_status()
with open(output_path, "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk) The cURL equivalent uses the --output flag to write directly to disk:
curl -X GET "https://na1.fusion.foxit.com/pdf-services/api/documents/<result_document_id>/download" \
-H "client_id: $CLIENT_ID" \
-H "client_secret: $CLIENT_SECRET" \
--output output.pdf To verify the output, check response.headers.get("Content-Type") for application/pdf, or inspect the first four bytes of the written file for the %PDF magic bytes if your pipeline requires format validation.
Error Handling for Production
The Foxit PDF Services API documentation covers 400, 404, 413, and 500 across the four endpoints. The 401 appears on authentication failures as a practical case even though it’s absent from the documented example responses. Each status code points to a specific root cause with a concrete recovery path:
- 400: malformed request body or unsupported file type. Validate the input file path and extension before calling
upload_doc(). - 401: credential misconfiguration. Verify that CLIENT_ID and CLIENT_SECRET are exported in your shell and that the header names are lowercase
client_idandclient_secret. - 404: the
documentIdhas expired. The server deletes uploaded files after 24 hours, so the convert and download endpoints return 404 for anydocumentIdpast that window. Re-upload the source file and restart from the upload step. An expired or unknowntaskIdon the poll endpoint behaves differently: it returns HTTP 200 withstatus: "FAILED"and anerrorobject whosemessagereads"task is not exist". The poll loop’sFAILEDbranch already catches that case. - 413: file exceeds the 100 MB upload cap. Pre-check with
os.path.getsize()before uploading, or split the document. - 500: transient server error. Apply exponential backoff with a ceiling of 3 retries (wait times of 1s, 2s, and 4s).
def call_with_retry(fn, *args, max_retries: int = 3, **kwargs):
for attempt in range(max_retries + 1):
try:
return fn(*args, **kwargs)
except requests.HTTPError as e:
code = e.response.status_code
if code == 400:
raise ValueError(
"Bad request. Confirm the input is a supported Word format."
) from e
if code == 401:
raise PermissionError(
"Authentication failed. Check CLIENT_ID and CLIENT_SECRET env vars."
) from e
if code == 404:
raise FileNotFoundError(
"Document or task expired (24h TTL). Re-upload and retry."
) from e
if code == 413:
raise OverflowError(
"File too large. The upload cap is 100 MB."
) from e
if code == 500 and attempt < max_retries:
wait = 2 ** attempt # 1s, 2s, 4s
print(f"Server error. Retrying in {wait}s ({attempt + 1}/{max_retries})")
import time
time.sleep(wait)
continue
raise Pipeline authors should treat documentId values as ephemeral: each one expires with its source file after 24 hours, so pipeline code that caches documentId values between sessions will see 404s on every convert call, and re-uploading is always the correct recovery path.
The Complete Script
Set your environment variables, then run python convert.py input.docx output.pdf:
import os
import json
import time
import sys
import requests
CLIENT_ID = os.environ.get("CLIENT_ID")
CLIENT_SECRET = os.environ.get("CLIENT_SECRET")
BASE_URL = "https://na1.fusion.foxit.com"
def call_with_retry(fn, *args, max_retries: int = 3, **kwargs):
for attempt in range(max_retries + 1):
try:
return fn(*args, **kwargs)
except requests.HTTPError as e:
code = e.response.status_code
if code == 400:
raise ValueError(
"Bad request. Confirm the input is a supported Word format."
) from e
if code == 401:
raise PermissionError(
"Authentication failed. Check CLIENT_ID and CLIENT_SECRET env vars."
) from e
if code == 404:
raise FileNotFoundError(
"Document or task expired (24h TTL). Re-upload and retry."
) from e
if code == 413:
raise OverflowError(
"File too large. The upload cap is 100 MB."
) from e
if code == 500 and attempt < max_retries:
wait = 2 ** attempt # 1s, 2s, 4s
print(f"Server error. Retrying in {wait}s ({attempt + 1}/{max_retries})")
time.sleep(wait)
continue
raise
def upload_doc(file_path: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/upload"
headers = {"client_id": CLIENT_ID, "client_secret": CLIENT_SECRET}
with open(file_path, "rb") as f:
files = {"file": (os.path.basename(file_path), f)}
r = requests.post(url, headers=headers, files=files)
r.raise_for_status()
return r.json()["documentId"]
def convert_to_pdf(document_id: str) -> str:
url = f"{BASE_URL}/pdf-services/api/documents/create/pdf-from-word"
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
"Content-Type": "application/json",
}
r = requests.post(url, headers=headers, data=json.dumps({"documentId": document_id}))
r.raise_for_status()
return r.json()["taskId"]
def poll_task(task_id: str, max_attempts: int = 30) -> str:
url = f"{BASE_URL}/pdf-services/api/tasks/{task_id}"
headers = {"client_id": CLIENT_ID, "client_secret": CLIENT_SECRET}
for attempt in range(max_attempts):
r = requests.get(url, headers=headers)
r.raise_for_status()
data = r.json()
status = data.get("status")
print(f"[{attempt + 1}/{max_attempts}] status={status}, progress={data.get('progress', 0)}%")
if status == "COMPLETED":
return data["resultDocumentId"]
if status == "FAILED":
raise RuntimeError(f"Conversion failed: {data.get('error')}")
time.sleep(2)
raise TimeoutError(f"Task {task_id} did not complete after {max_attempts} attempts")
def download_result(result_document_id: str, output_path: str) -> None:
url = f"{BASE_URL}/pdf-services/api/documents/{result_document_id}/download"
headers = {"client_id": CLIENT_ID, "client_secret": CLIENT_SECRET}
with requests.get(url, headers=headers, stream=True) as r:
r.raise_for_status()
with open(output_path, "wb") as f:
for chunk in r.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
def convert_docx_to_pdf(input_path: str, output_path: str) -> None:
print(f"Uploading {input_path}...")
document_id = call_with_retry(upload_doc, input_path)
print(f"Uploaded. documentId={document_id}")
print("Initiating conversion...")
task_id = call_with_retry(convert_to_pdf, document_id)
print(f"Queued. taskId={task_id}")
print("Polling for completion...")
result_document_id = call_with_retry(poll_task, task_id)
print(f"Completed. resultDocumentId={result_document_id}")
print(f"Downloading to {output_path}...")
call_with_retry(download_result, result_document_id, output_path)
print("Done.")
if __name__ == "__main__":
if len(sys.argv) != 3:
print("Usage: python convert.py <input.docx> <output.pdf>")
sys.exit(1)
convert_docx_to_pdf(sys.argv[1], sys.argv[2]) The Foxit PDF Services API also supports merging, compression, linearization, and OCR through additional endpoints. All of them share the same host and header-based auth pattern, so the functions you’ve built here extend naturally as your pipeline grows.
Create your free Foxit developer account and run your first conversion in under five minutes, with no credit card required at signup.
DOCX to PDF API FAQ
Does the Foxit PDF Services API support formats other than .docx?
Yes. The /pdf-services/api/documents/create/pdf-from-word endpoint accepts .doc, .docx, .rtf, .dot, .dotx, .docm, .dotm, and .wpd. The same four-step flow applies for all of them.
How long are uploaded files retained?
Uploaded documents are automatically deleted after 24 hours. Treat documentId values as ephemeral and re-upload whenever you need to convert a file after that window.
What happens if I poll the task endpoint faster than every 2 seconds?
Faster polling consumes rate limit budget without affecting conversion speed. The server determines conversion time based on document complexity and queue load, so polling intervals below 2 seconds add no throughput benefit.
Can I run multiple DOCX-to-PDF conversions in parallel?
Yes. Each upload returns an independent documentId and each conversion returns an independent taskId. Run concurrent conversions by launching multiple threads or async tasks, with each one tracking its own taskId. Python’s concurrent.futures.ThreadPoolExecutor is a straightforward way to manage this.
Where do I get my CLIENT_ID and CLIENT_SECRET?
From the Foxit Developer Portal dashboard, under the default application created at signup. Both values are available immediately after account creation.
Does the API require an OAuth token exchange?
The API authenticates through named request headers. Pass client_id and client_secret directly on every request, and the server reads those credentials on each call.
PDF Translation with Verifiable Quality: Build a Confidence-Scored Pipeline with Foxit API and Straker.ai

Most machine translation tools hand back a translated PDF with no signal about which parts to trust — a real problem for contracts, medical forms, and regulatory filings. This guide shows how to build a pipeline that scores every segment before the final render, using Foxit for structural extraction and layout-preserving rendering and Straker.ai for translation plus per-segment quality scoring.
Most machine translation tools give you a translated file and nothing else. They do not tell you which parts are correct and which parts are wrong. For a simple blog post, that is fine. For a contract, a medical form, or a legal notice, it is a real problem. A bad translation can sit in the final PDF for days before anyone notices, often only after the document has already been signed or sent.
Teams today are translating more documents, into more languages, and faster than ever. Legal, finance, healthcare, HR, and insurance teams all deal with PDFs where one wrong word can cause a lot of damage: a broken contract, a failed audit, or even a safety issue. Most translation tools were not built to catch these mistakes. They just move text from one language to another. When quality checks happen at all, they usually mean a person reading the final PDF line by line and hoping they spot the errors.
This article shows how to build a better setup. You will learn how to build a PDF translation pipeline that gives every segment a quality score before the final PDF is created. Instead of hoping the translation is right, the pipeline tells you which parts to trust, which parts to review, and which parts to send back to a human translator. All of this happens automatically on every run.
Architecture at a Glance
Before going deeper, it helps to see the full pipeline in one picture. The diagram below traces a source PDF through every stage: extract, translate, score, route, and render. Each box is a single responsibility handled by a single service, with the routing layer acting as the glue you control.

The pipeline has two external services:
- Foxit PDF Translation API handles anything PDF-specific. It pulls the structured text out of the source document with element IDs attached, then renders the final PDF back in the original layout (multi-column text, tables, font substitution, image positions) using the approved translations.
- Straker AI translates each source segment AND scores the translation in the same request. It returns the target text, a numeric score on a 0.0 to 1.0 scale, and a categorical label (
best,good,acceptable,bad) for every element ID. This step is pluggable, so you can swap Straker for DeepL, Google Cloud Translation, AWS Translate, or an in-house NMT if you already have a contract with one of them. The contract between this step and the rest of the pipeline is a flat dict of element IDs to translated text plus per-segment scores.
and one piece of code you own:
- Routing layer is your business logic. It reads the score, decides whether the segment auto-accepts, flags for human review, or escalates to a translator, and then hands the approved set to Foxit’s render call.
With the shape of the pipeline on the table, the rest of the article works through each piece in order, starting with why per-segment quality scoring is worth the integration effort in the first place.
The Quality Gap
You ship a translated PDF to a legal team. Three days later, compliance flags a clause in the German version. The term “indemnification” was rendered as “Entschädigung” (compensation) rather than “Freistellung” (hold harmless). Your MT pipeline returned a 200 status. Nobody’s alerting on that delta.
Raw machine translation output carries no quality signal by default. Every segment comes back translated, and your pipeline treats them identically regardless of whether the model was confident or guessing. For marketing copy that’s an acceptable tradeoff, but for a loan covenant, a clinical trial protocol, or a regulatory filing, a 95%-accurate translation can still be contractually or legally dangerous because the 5% failure may concentrate precisely in the high-stakes clauses.
A confidence score, in the translation QA context, is a per-segment numeric signal from a verification engine. It tells you how reliable each translated unit is on a scale your system can act on programmatically. High-confidence segments auto-accept, medium-confidence ones queue for post-edit review, and low-confidence segments escalate directly to a human translator before they ever reach the final document.
The compound problem for PDFs specifically is that most translation pipelines strip document structure before the MT engine even sees the text. The extraction step flattens multi-column layouts, collapses table cells, and drops font metadata. By the time you get a translated output, you’ve lost both layout fidelity and any quality signal. The rendered PDF looks wrong and you have no programmatic way to know which segments caused it.
Foxit’s PDF Translation Trial API extracts structured text from a source PDF with element IDs preserved, so the layout blueprint travels alongside the text through the entire workflow. You hand the source segments to Straker AI, which returns the translated text plus a per-segment numeric score and a quality label in a single call. (If you already run DeepL, Google Cloud Translation, AWS Translate, or an in-house NMT Engine, you can drop it in at this step without changing the rest of the pipeline.) Your routing logic decides which segments pass, which get flagged, and which escalate to human review. Foxit’s render endpoint then re-assembles the PDF in the original layout using the accepted translations, giving you a layout-preserved translated PDF with a documentable quality trail attached to every segment.
How the Pipeline Works
Foxit and Straker are two independent APIs that you wire together. Foxit owns PDF structure, extracting structured text keyed by element ID and re-rendering the final PDF in the original layout. Straker AI handles translation and per-segment quality scoring in a single request, returning the translated text alongside a numeric score and a quality label. You own the routing decision that sits between the scores and the render call.
The pipeline runs in seven steps:

Foxit covers steps 1-3 and 6-7 (PDF structure and rendering). Straker AI covers step 4, producing translations and per-segment quality scores in one round-trip. Step 5 is your business logic.
The Foxit PDF Translation API defines steps 2, 3, and 6. The upload and download calls use the general PDF Services endpoints. Straker AI is a separate API at https://api-verify.straker.ai. You submit XLF 1.2 files containing source segments and Straker returns the translated target_text per segment plus a numeric score (0.0 to 1.0) and a quality label (best, good, acceptable, bad). Because Foxit’s ExtractedText.json is a flat { "elementId": "text" } map, and XLF trans-unit IDs round-trip through Straker’s external_id field unchanged, the element IDs Foxit emits are the same IDs that come back with translations and scores attached. That alignment is what makes programmatic routing possible.
One clarification for readers who’ve seen the Foxit-Straker partnership announcement: that partnership covers Foxit eSignature Services, enabling end users to translate and sign documents in the eSign product. That’s an end-user feature. The PDF Translation Trial API used here is a separate developer surface. Its OpenAPI spec (v2.2.0) contains zero Straker references, and the preprocess-pdf documentation explicitly instructs developers to “translate the text in ExtractedText.json using your preferred translation tool.” You wire the two APIs together manually. This tutorial uses Straker AI as the default translation engine because it produces translations and quality scores in the same call, but you can substitute DeepL, Google Cloud Translation, AWS Translate, or your own NMT at step 4 without changing the Foxit calls.
Credentials and Setup
Get your Foxit credentials at app.developer-api.foxit.com/pricing. The free Developer plan gives you 20 AI credits per month with no credit card and no sales call required. Once you’ve signed in, your Client ID and Client Secret appear in the developer dashboard. Every Foxit API call requires both in the request headers as client_id and client_secret (lowercase snake_case). Export them in your shell as FOXIT_CLIENT_ID and FOXIT_CLIENT_SECRET so the code below reads them from the environment rather than hard-coding secrets.
For Straker, sign up at straker.ai/ai-platform/verify for API access. Straker issues a UUID-style API token that you send as a bearer token on every call (Authorization: Bearer <your-token>). The API lives at https://api-verify.straker.ai and its full reference is published at api-verify.straker.ai/docs. Export your token as STRAKER_API_KEY for the code below. You can confirm the token works and check your balance with a quick GET /user/balance. Both services offer trial access, so you can build and test the full pipeline before any procurement conversation.
Before you finalize your language matrix, check both APIs for supported languages. Foxit’s render endpoint accepts 23 target language codes (en, zh, zh_tw, fr, de, es, it, pt, nl, ja, ko, th, vi, hi, ru, ar, tr, pl, sv, no, nb, da, and fi). Straker AI identifies languages by UUID rather than ISO code. You fetch the full list with GET /languages and look up the UUID for your target (for example, 917FF728-0725-A033-1278-33025F49CA40 is French (France), 917FF7D8-9107-0BF8-97EE-065C20F453DE is German). The intersection of the two sets determines your production language coverage.
If you already have a contract with DeepL, Google Cloud Translation, AWS Translate, or an in-house NMT service, you can swap that engine in at step 4. The pipeline contract upstream (Foxit element IDs mapped to source strings) and downstream (a dict of {element_id: {score, quality, target_text}} feeding the router) does not change. The code below uses Straker AI by default because the same API returns the translation and the quality signal in one call.
Building the PDF Translation Pipeline
The complete seven-step pipeline runs in Python using requests, json, zipfile, os, and the standard-library xml.etree.ElementTree for building XLF. The first snippet covers Foxit steps 1-3 (upload, structural extraction, and preprocessing).
import requests
import json
import zipfile
import io
import time
FOXIT_BASE = "https://na1.fusion.foxit.com/pdf-services/api"
HEADERS = {
"client_id": "YOUR_CLIENT_ID",
"client_secret": "YOUR_CLIENT_SECRET"
}
def poll_task(task_id: str) -> dict:
"""Poll GET /tasks/{task_id} until COMPLETED or FAILED."""
while True:
r = requests.get(f"{FOXIT_BASE}/tasks/{task_id}", headers=HEADERS)
r.raise_for_status()
data = r.json()
status = data.get("status")
if status == "COMPLETED":
return data
if status == "FAILED":
raise RuntimeError(f"Task {task_id} failed: {data.get('error')}")
# PENDING or IN_PROGRESS: wait and retry
time.sleep(3)
# Step 1: Upload source PDF
with open("source.pdf", "rb") as f:
upload_resp = requests.post(
f"{FOXIT_BASE}/documents/upload",
headers=HEADERS,
files={"file": ("source.pdf", f, "application/pdf")}
)
upload_resp.raise_for_status()
source_document_id = upload_resp.json()["documentId"]
# Step 2: Structural Extract (async - must complete before preprocess)
extract_resp = requests.post(
f"{FOXIT_BASE}/documents/pdf-structural-extract",
headers=HEADERS,
json={"documentId": source_document_id}
)
extract_resp.raise_for_status() # 202 Accepted
extract_task_id = extract_resp.json()["taskId"]
extract_result = poll_task(extract_task_id)
extracted_doc_id = extract_result["resultDocumentId"]
# Step 3: Preprocess (synchronous - returns 200, no polling needed)
preprocess_resp = requests.post(
f"{FOXIT_BASE}/documents/translation/preprocess-pdf",
headers=HEADERS,
json={"documentId": extracted_doc_id}
)
# Errors from preprocess-pdf per the Foxit spec:
# 400 VALIDATION_ERROR - "Document ID is required"
# 500 INTERNAL_SERVER_ERROR - "Failed to preprocess document"
preprocess_resp.raise_for_status()
preprocess_result_id = preprocess_resp.json()["resultDocumentId"]
# Download the ZIP containing ExtractedText.json and StructureInfo.json
zip_resp = requests.get(
f"{FOXIT_BASE}/documents/{preprocess_result_id}/download",
headers=HEADERS
)
zip_resp.raise_for_status()
with zipfile.ZipFile(io.BytesIO(zip_resp.content)) as zf:
extracted_text = json.loads(zf.read("ExtractedText.json"))
# StructureInfo.json: do not modify - the render step requires it untouched
# structure_info = json.loads(zf.read("StructureInfo.json"))
# extracted_text is now {"elementId1": "original text", "elementId2": "original text", ...} The preprocess step is synchronous, which means you get a 200 OK directly with the resultDocumentId. No polling required. The ZIP it produces contains two files: ExtractedText.json maps every element ID to its original text, and StructureInfo.json carries the full layout blueprint (bounding boxes, font metadata, column positions). You pass StructureInfo.json to the render step unmodified. Modifying it breaks the render because it’s the mechanism that makes layout preservation possible.
The second snippet covers steps 4-7, calling Straker AI to translate and score every segment in one round-trip, routing by score, rendering the translated PDF, and downloading the result. Straker’s AI Translation and Quality Evaluation workflow accepts a source-only XLF and returns a translated target_text per segment alongside the numeric score and the quality label, so the same response feeds both the translation choice and the routing decision.
import xml.etree.ElementTree as ET
STRAKER_BASE = "https://api-verify.straker.ai"
STRAKER_TOKEN = "STRAKER_API_KEY"
STRAKER_HEADERS = {"Authorization": f"Bearer {STRAKER_TOKEN}"}
# Straker identifies languages by UUID. Look these up once via GET /languages
# and cache them. Full list: https://api-verify.straker.ai/languages
STRAKER_LANG_FRENCH = "917FF728-0725-A033-1278-33025F49CA40"
STRAKER_LANG_GERMAN = "917FF7D8-9107-0BF8-97EE-065C20F453DE"
# Workflow UUID for "AI Translation and Quality Evaluation". Fetch the full
# list of workflows once via GET /workflow and cache the UUID for the one you
# want; this workflow produces both the translation and the per-segment score.
STRAKER_WORKFLOW_AI_TRANSLATE_AND_EVAL = "390b47a9-d5dc-46ae-92e2-56c43d128c44"
def build_xlf_1_2_source_only(source_lang: str, target_lang: str,
sources: dict) -> bytes:
"""
Build a minimal XLF 1.2 document with source segments and empty targets.
trans-unit/@id preserves Foxit's element IDs; Straker surfaces the same
value as `external_id` on the segments it returns, so the keys round-trip.
"""
ns = "urn:oasis:names:tc:xliff:document:1.2"
ET.register_namespace("", ns)
xliff = ET.Element(f"{{{ns}}}xliff", {"version": "1.2"})
file_el = ET.SubElement(xliff, f"{{{ns}}}file", {
"source-language": source_lang,
"target-language": target_lang,
"datatype": "plaintext",
"original": "foxit-extract",
})
body = ET.SubElement(file_el, f"{{{ns}}}body")
for element_id, source_text in sources.items():
unit = ET.SubElement(body, f"{{{ns}}}trans-unit", {"id": element_id})
ET.SubElement(unit, f"{{{ns}}}source").text = source_text
ET.SubElement(unit, f"{{{ns}}}target") # empty - Straker fills it in
return b'<?xml version="1.0" encoding="UTF-8"?>\n' + ET.tostring(xliff, encoding="utf-8")
# Step 4: Translate and score every segment with Straker AI in one call.
def translate_and_score_with_straker(sources: dict, source_lang_code: str,
target_lang_uuid: str) -> dict:
"""
Submit source-only XLF to Straker's AI Translation + Quality Evaluation
workflow. Returns a dict keyed by Foxit element ID ->
{"score": float|None, "quality": str, "target_text": str}.
"""
xlf_bytes = build_xlf_1_2_source_only(source_lang_code, "fr", sources)
# 4a. Create the project on the AI Translation + Quality Evaluation
# workflow. confirmation_required=false commits the token cost
# immediately; set to true to review cost and call POST /project/confirm
# before processing begins.
create_resp = requests.post(
f"{STRAKER_BASE}/project",
headers=STRAKER_HEADERS,
files={"files": ("segments.xlf", xlf_bytes, "application/xliff+xml")},
data={
"languages": target_lang_uuid,
"title": "Foxit PDF translation batch",
"workflow_id": STRAKER_WORKFLOW_AI_TRANSLATE_AND_EVAL,
"confirmation_required": "false",
},
)
create_resp.raise_for_status()
project_id = create_resp.json()["project_id"]
# 4b. Poll the project until it reports COMPLETED.
while True:
status_resp = requests.get(
f"{STRAKER_BASE}/project/{project_id}", headers=STRAKER_HEADERS
)
status_resp.raise_for_status()
project = status_resp.json()["data"]
if project["status"] == "COMPLETED":
break
if project["status"] in ("FAILED", "PROCESSING_FAILED", "CANCELED"):
raise RuntimeError(f"Straker project {project_id} failed")
time.sleep(3)
# 4c. Fetch the per-segment translations + scores. file_uuid is returned
# in the project payload.
file_uuid = project["source_files"][0]["file_uuid"]
seg_resp = requests.get(
f"{STRAKER_BASE}/project/{project_id}/segments/{file_uuid}/{target_lang_uuid}",
headers=STRAKER_HEADERS,
)
seg_resp.raise_for_status()
results = {}
for seg in seg_resp.json()["segments"]:
element_id = seg["external_id"] # matches the Foxit key we packed into XLF
t = seg["translation"]
results[element_id] = {
"score": t["score"], # float 0.0 to 1.0, or None
"quality": t["quality"], # "best" | "good" | "acceptable" | "bad"
"target_text": t["target_text"], # Straker's translation
}
return results
scored = translate_and_score_with_straker(
extracted_text,
source_lang_code="en",
target_lang_uuid=STRAKER_LANG_FRENCH,
)
# Step 5: Route by score and quality label (developer-controlled business logic).
HIGH_THRESHOLD = 0.85
LOW_THRESHOLD = 0.65
accepted = {}
flagged_for_review = {}
rejected = {}
for element_id, verdict in scored.items():
score = verdict["score"] or 0.0
if verdict["quality"] == "best" or score >= HIGH_THRESHOLD:
accepted[element_id] = verdict["target_text"]
elif verdict["quality"] == "bad" or score < LOW_THRESHOLD:
rejected[element_id] = {"original": extracted_text[element_id],
"score": score, "quality": verdict["quality"]}
else:
flagged_for_review[element_id] = {"translation": verdict["target_text"],
"score": score, "quality": verdict["quality"]}
# Build the render payload. Foxit's render expects every key from the original
# ExtractedText.json. Accepted segments use the scored translation; flagged and
# rejected segments fall back to the original source text so the layout is not
# broken by missing keys. In production, replace the fallback with human-
# reviewed text once it is available, or hold the render step until review
# completes.
render_payload = {}
for element_id, original_text in extracted_text.items():
if element_id in accepted:
render_payload[element_id] = accepted[element_id]
else:
render_payload[element_id] = original_text
# Step 6: Render (async)
# translatedFile is the modified ExtractedText.json with translated values, same keys
translated_json_bytes = json.dumps(render_payload).encode("utf-8")
render_resp = requests.post(
f"{FOXIT_BASE}/documents/translation/render-pdf",
headers=HEADERS,
data={
"sourceDocumentId": source_document_id,
"preprocessResultDocumentId": preprocess_result_id,
"targetLanguage": "fr"
# Optional: "pageRangeStart": 1, "pageRangeEnd": 10
},
files={"translatedFile": ("ExtractedText.json", translated_json_bytes, "application/json")}
)
# Errors from render-pdf per the Foxit spec:
# 400 VALIDATION_ERROR - "Either translatedFile or translatedTextDocumentId must be provided"
# 400 VALIDATION_ERROR - "Unsupported target language: xx"
# 500 RENDER_START_FAILED - "Failed to start render: service unavailable"
render_resp.raise_for_status()
render_task_id = render_resp.json()["taskId"]
render_result = poll_task(render_task_id)
output_doc_id = render_result["resultDocumentId"]
# Step 7: Download translated PDF
pdf_resp = requests.get(
f"{FOXIT_BASE}/documents/{output_doc_id}/download",
headers=HEADERS
)
pdf_resp.raise_for_status()
with open("translated_output.pdf", "wb") as f:
f.write(pdf_resp.content)
print(f"Done. Accepted: {len(accepted)}, Flagged: {len(flagged_for_review)}, Rejected: {len(rejected)}") The render call is multipart/form-data. You pass sourceDocumentId (the original PDF’s document ID from step 1), preprocessResultDocumentId (from step 3), targetLanguage (one of the 23 supported codes), and translatedFile (the modified ExtractedText.json with translated values and original keys). The alternative is uploading the translated JSON first via the upload endpoint and passing its ID as translatedTextDocumentId instead. At least one of the two must be present, or you’ll get a 400 VALIDATION_ERROR.
The render operation is asynchronous. It returns 202 Accepted immediately with a taskId, and the actual rendering runs in the background on Foxit’s side. You must poll GET /tasks/{taskId} on a fixed interval, every 3 seconds is the recommended cadence, until the status flips to COMPLETED before you try to download the output. Skipping the poll, or treating the initial 202 response as if it were a finished render, will cause the program to crash and interrupt the rest of the pipeline because the result document is not yet written when the task is still IN_PROGRESS. The poll_task helper from the first snippet already implements this loop with a 3-second time.sleep between checks and surfaces a FAILED status as a RuntimeError, so reuse it here rather than reading render_resp.json() directly. The same polling discipline applies to the structural extract step (step 2), which is also asynchronous.
Scoring and Routing
Straker AI generates both the translation and the quality signal in this pipeline. Foxit’s responses carry document IDs and task statuses; the translation choice and the per-segment score are entirely Straker’s contribution.
Each segment in the /project/{id}/segments/{file_id}/{language_id} response carries three values you care about. target_text is Straker’s translation. score is a float between 0.0 and 1.0 (it may be null for segments where the model has no confidence signal). quality is a categorical label Straker assigns alongside the numeric score (best, good, acceptable, or bad). You can route on either signal, or combine them. The table below shows a combined policy calibrated for compliance-sensitive documents. These are starting points; your production system should calibrate per language pair and domain, since a French legal contract demands different thresholds than a Spanish marketing brochure.
| Straker verdict | Action | Rationale |
|---|---|---|
quality == "best" or score >= 0.85 | Auto-accept, include in render | High confidence output; suitable for fully automated workflows |
quality in ("good", "acceptable") or 0.65 - 0.84 | Flag segment by element ID for post-edit review | Medium confidence; a human reviewer checks the flagged segments before the final render runs |
quality == "bad" or score < 0.65 | Reject segment, escalate to human translator | Low confidence output; the model is unreliable for this segment |
The element ID key structure matters here. Foxit’s ExtractedText.json keys are packed into XLF trans-unit IDs, and Straker surfaces the same value in its response’s external_id field. That means every entry in your flagged_for_review dictionary carries enough information for a reviewer to open the source document, find the exact element by ID, and return an approved translation. You write the approved translation back into the same key, then trigger the render step. This produces a documentable audit trail. For every element ID in the output PDF, you can show the original text, Straker’s translation, the Straker score and quality label, and whether a human approved it. In regulated industries (finance, legal, healthcare), that’s the evidence your compliance team needs to sign off on an automated localization workflow, and it aligns with ISO 18587, the international standard for post-editing of machine translation output.
Straker AI can also route low-confidence output to expert reviewers automatically when configured through the Straker platform. Check straker.ai/ai-platform/verify for the workflow configuration options.
Layout Preservation
Foxit’s render step preserves multi-column text flow, embedded table cell structure, images at their original positions, headers and footers, and font substitution for target-language character sets. That means CJK scripts (Japanese, Chinese, Korean) render correctly with appropriate glyph substitution, and Arabic output renders right-to-left without manual post-processing.
StructureInfo.json is what makes this possible. When the preprocess step runs, it produces both the text map (which you hand to Straker) and the layout blueprint (which you hand back to Foxit unmodified at render time). The render engine maps translated text back to the original element positions using this blueprint, reflowing text within the same bounding boxes. Because the structure data travels alongside the text through the entire pipeline, Foxit never needs to reconstruct the layout from scratch.
Generic MT pipelines export raw text, losing all spatial relationships, translate it, then attempt to rebuild the PDF from nothing. Tables merge into continuous text, columns collapse to a single flow, and CJK font substitution fails because the rebuilding step has no record of what fonts were originally in use.
Limitations to Test
Text expansion is the first limitation worth stress-testing. English to German translation typically increases text length by 20-35%, and English to Arabic can run even longer. Foxit’s render engine handles reflow within bounding boxes, but extreme length changes in tight table cells or narrow columns may overflow. Test with your actual document types before you commit to a production deployment.
Complex layout edge cases are the second limitation. Overlapping text boxes, embedded SVG charts with text labels, and PDFs with non-standard encoding may produce imperfect renders. The structural extraction step covers standard PDF text elements well, but edge-case layouts require manual review of the rendered output before you sign off on the pipeline for a given document class.
Try It Now
Sign up for Foxit’s free Developer plan and a Straker AI account, grab credentials for both, and run the pipeline from the section above against a real document. An invoice, a multi-page contract, or a regulatory filing works well for testing because each has tables, mixed-column layouts, and high-stakes text segments.
After the render completes, verify four things in the output PDF:
- Tables retain cell structure
- Multi-column text flows correctly in the target language
- Images remain in their original positions
- Fonts render correctly for the target script
Cross-reference the confidence scores from Straker against the rendered segments to calibrate your production thresholds. You may find that legal terminology in German warrants a 0.90 auto-accept threshold while product description text in French is fine at 0.80.
The complete Foxit Translation Trial API reference covers the full parameter list and response schema for preprocess-pdf and render-pdf. The Foxit Structural Extraction Trial API reference documents the structural extract endpoint. Straker’s translation and scoring API documentation lives at straker.ai/ai-platform/verify.
Looking ahead, Straker’s dashboard lists a native Foxit integration as Coming Soon (no release date announced at the time of writing), described as a workflow to translate PDF contracts with Foxit, verify them with experts, and finalize them for signing. When it ships, it’s likely to compress several of the manual steps above into a single call. The underlying mechanics (structural extract, translation, per-segment scoring, routing, render) will remain the same logical stages, so the pipeline you build today stays a useful mental model for reasoning about the native version when it arrives.
For production-scale implementation patterns and how Straker’s translation and verification layer integrates into enterprise localization pipelines, register for the upcoming joint Foxit + Straker.ai webinar with Lee Konstanty from Straker. Get your Foxit API credentials | Get started with Straker AI
PDF Translation API FAQ
What is a PDF translation API with confidence scoring?
A PDF translation API with confidence scoring is a service that translates PDF documents and returns a per-segment quality signal alongside each translation. Instead of handing back a single translated file, the API tells you which segments are high-confidence (safe to auto-accept), which are medium-confidence (queue for human review), and which are low-confidence (escalate to a translator). This pipeline combines Foxit’s PDF Translation Trial API for structural extraction and layout-preserving rendering with Straker.ai for translation and scoring in a single call.
How does the Foxit and Straker.ai PDF translation pipeline work?
The pipeline runs in seven steps: upload the source PDF to Foxit, run structural extraction to get element-ID-keyed text, preprocess to produce ExtractedText.json and StructureInfo.json, send segments to Straker AI’s “AI Translation and Quality Evaluation” workflow which returns translated text plus a 0.0–1.0 score and a quality label, route each segment programmatically by score, then call Foxit’s render endpoint to rebuild the PDF in the original layout. Foxit owns PDF structure, Straker owns translation and scoring, and your code owns the routing decision.
Why do PDFs need per-segment translation quality scores?
For marketing copy, raw machine translation output is usually fine. For contracts, medical forms, clinical trial protocols, or regulatory filings, a 95%-accurate translation can still be legally dangerous because the 5% failure may land on a high-stakes clause — like “indemnification” rendered as “Entschädigung” (compensation) instead of “Freistellung” (hold harmless). Per-segment confidence scores let you route low-confidence segments to human reviewers before they reach the final document, producing the audit trail compliance teams need under standards like ISO 18587.
Can I use DeepL, Google Translate, or AWS Translate instead of Straker.ai?
Yes. The translation step is pluggable. The contract upstream — Foxit element IDs mapped to source strings — and downstream — a dict of element IDs to translated text feeding the render call — does not change if you swap the engine. DeepL, Google Cloud Translation, AWS Translate, or an in-house NMT engine all work. The trade-off is that Straker AI returns translation plus quality score in one call, while other engines require a separate verification step if you want confidence signals.
How does Foxit preserve PDF layout during translation?
Foxit’s preprocess step produces two files: ExtractedText.json with element-ID-keyed text, and StructureInfo.json with the full layout blueprint (bounding boxes, font metadata, column positions, image locations). You modify only ExtractedText.json with translations and pass StructureInfo.json to the render endpoint untouched. The render engine reflows translated text within the original bounding boxes, handles font substitution for CJK and Arabic scripts, and preserves multi-column layouts, tables, and image positions — without rebuilding the PDF from scratch.
What target languages does the Foxit PDF Translation API support?
Foxit’s render endpoint accepts 23 target language codes: en, zh, zh_tw, fr, de, es, it, pt, nl, ja, ko, th, vi, hi, ru, ar, tr, pl, sv, no, nb, da, and fi. Straker AI identifies languages by UUID rather than ISO code, fetched via GET /languages. Your production language coverage is the intersection of both sets — check both APIs before finalizing your language matrix.
How do I set confidence score thresholds for auto-accept versus human review?
A reasonable starting policy for compliance-sensitive documents: auto-accept segments with quality == “best” or score >= 0.85, flag for post-edit review at 0.65–0.84 or quality in (“good”, “acceptable”), and reject for human translation at score < 0.65 or quality == “bad”. These are starting points — calibrate per language pair and domain. A French legal contract may warrant a 0.90 auto-accept threshold while a Spanish marketing brochure is fine at 0.80. Run the pipeline against a representative sample of your real documents and tune from there.
Extract Anything from Any PDF: Inside Foxit’s Advanced Extraction Engine

Basic PDF extraction libraries break on scanned documents, complex tables, and form fields, leaving downstream pipelines starved of clean data. Foxit’s PDF Structural Extraction API combines OCR, layout recognition, and AI parsing to return all twelve PDF element types as structured JSON, ready for RAG, BI, and CRM workflows.
Your PDF extraction pipeline passes unit tests against the sample invoices you built it on. Then production arrives and you’re looking at 47% garbled output on the Q4 contract batch because half those documents are scanned TIFFs wrapped in a PDF envelope, and your extraction library has no concept of what an image-only page actually is.
The failure modes are specific. PyMuPDF’s get_text() returns empty strings on scanned PDFs because it reads content streams directly, and image-only pages carry no text stream. pdfplumber’s table detection merges rows when column widths span non-uniform grids, which is standard in any financial statement that mixes summary and line-item rows on the same page. Embedded images containing meaningful text (stamped signatures, engineering drawing annotations, letterhead logos) get silently dropped. The library extracts coordinates for the XObject reference but does nothing with the raster data inside. Form fields built on non-standard annotation types (AcroForms using widget annotations with custom action streams) lose their values entirely when you serialize to text.
The architectural distinction that creates this problem is the difference between content serialization and semantic extraction. A PDF converter reads a content stream and writes out whatever character sequences it finds in rendering order. An extraction engine understands the spatial relationships between those character sequences: that two columns of text at x=72 and x=320 are parallel body copy, that the row at y=210 belongs to the table starting at y=180, that the text block repeating on every page is a header carrying lower retrieval weight in a RAG index. Output that lacks spatial and semantic classification looks correct on screen but breaks every downstream consumer that depends on structure.
BI dashboards require numbers tied to the right row labels. AI ingestion pipelines require heading hierarchy to chunk accurately. CRMs require form field values extracted from AcroForm widget dictionaries, delivered with field names intact. The delta between what basic extraction libraries return and what those systems can actually consume is where document pipeline engineering hours accumulate.
How Foxit’s PDF Structural Extraction Engine Works Under the Hood
Foxit exposes this capability as the PDF Structural Extraction (Trial) endpoint inside the PDF Services API (POST /pdf-services/api/documents/pdf-structural-extract). Trial status means the schema is versioned at v1.0.7 and may evolve, but the contract is stable enough to build against today, and the endpoint runs against the production base URL at developer-api.foxit.com.
The engine runs three coordinated layers. The OCR layer operates on rasterized page content, recognizing characters from image-based PDFs and scanned documents across 200+ languages. The layout recognition layer applies spatial analysis to identify column boundaries, reading order, table cell boundaries, figure regions, and header/footer zones. The AI-based parsing layer classifies extracted objects semantically, resolving ambiguous blocks (a text run that spans two layout columns, or a figure caption that reads syntactically like a section heading) into typed elements.
All three layers run inside Foxit’s core PDF engine, which powers 700 million+ users across 20+ years of production deployments. That engine has native awareness of PDF internal structures: content streams, XObject dictionaries, AcroForm field trees, and annotation layers. The OCR layer operates on the same internal page representation the rendering engine uses, so it handles annotated PDFs where text overlaps image regions, and form fields where the visual display and stored value diverge.
The same Structural Extraction endpoint is also Step 1 of Foxit’s PDF Translation (Trial) workflow, which signals that the extraction output is structured enough to backbone a full rewrite-and-rerender pipeline.
NVIDIA’s July 2025 NeMo Retriever research on PDF extraction showed that specialized OCR-based pipelines outperform general-purpose vision-language models on retrieval recall and throughput for complex elements including tables, charts, and infographics. VLMs produce plausible-looking output on clean documents but degrade on exactly the edge cases (multi-column scans, mixed-content pages, annotated overlays) that a specialized pipeline handles systematically.
The Full Object Map: All 12 Extractable PDF Element Types
The Structural Extraction schema v1.0.7 defines twelve element types in the type enum: title, head, paragraph, table, image, headerFooter, form, hyperlink, footnote, sidebar, annotation, and formula.
The API exposes no per-object filter parameters. The only request body fields are documentId (required) and password (optional, for protected PDFs). The engine extracts the full element graph and returns everything in one asynchronous round-trip. You filter client-side on the returned JSON. The design is correct for the workload because partial extraction would require re-running layout recognition per request, costing more compute than transmitting the full element set in a single ZIP.
The result is a ZIP archive. At minimum it contains StructureInfo.json, whose top-level analyzeResult object holds version, pages, elements, and info. Documents that contain figures or tables also produce additional binary files (image renditions and table renditions) alongside the JSON, referenced from individual elements so the JSON payload stays manageable on large documents.
Each element in the document-wide flat elements array carries its own id, type, content, region (with page and an 8-point boundingBox polygon), and score confidence value. A table element adds its cell grid. A form element adds field data. An image element points to its binary file in the ZIP. Because title, head, and paragraph elements appear in document reading order in the elements array, they chunk cleanly on semantically correct boundaries, which is what a RAG index needs to return complete, coherent passages.
Each type maps directly to a downstream use case: table feeds financial reporting pipelines, form drives automated CRM data entry, image routes to computer vision workflows or document archives, annotation builds compliance audit trails, and head combined with paragraph elements in reading order feeds RAG ingestion.
API Walkthrough: The Four-Step Async PDF Extraction Flow
There’s no synchronous path. You upload, get a task ID, poll until completion, then download the result ZIP. Every request carries two headers: client_id and client_secret (lowercase snake_case, as specified in the API spec’s security schemes). Both come from the Developer Portal’s default application. Pass them as named HTTP headers on every request and do not use Authorization: Bearer.
The four-step sequence runs as follows:
The four-step sequence diagram uses two headers on every request: client_id and client_secret. Create a free developer account at account.foxit.com/site/sign-up (no credit card required, no sales call). Once you’re in, the credentials live under the default application in the Developer Portal. Copy the Client ID and Client Secret pair and treat them like any other API secret. Pass them as named HTTP headers on every call (lowercase snake_case, not Authorization: Bearer).
Step 1: Upload the PDF to
POST /pdf-services/api/documents/uploadasmultipart/form-datawith the file under field namefile. The 100MB ceiling is enforced with a413and error codeMAX_UPLOAD_SIZE_EXCEEDED. The response body returns{ "documentId": "doc_abc123" }.Step 2: Starts extraction with
POST /pdf-services/api/documents/pdf-structural-extract, passing{ "documentId": "doc_abc123" }. Add a"password"field for protected PDFs. The response is202 Acceptedwith{ "taskId": "task_xyz789" }.Step 3: Polls
GET /pdf-services/api/tasks/{task-id}. TheTaskResponsecarriestaskId,status,progress(0-100 integer),resultDocumentId, and an optionalerrorobject. Thestatusenum values arePENDING,IN_PROGRESS,COMPLETED, andFAILED. Portal narrative copy occasionally uses “PROCESSING,” but the schema enum value isIN_PROGRESS. Match your code against the enum. Poll untilCOMPLETEDand captureresultDocumentId.Step 4: Downloads with
GET /pdf-services/api/documents/{resultDocumentId}/download, which streams the ZIP archive. The optionalfilenamequery parameter overrides the default filename.
The complete cURL sequence for all four steps:
# Step 1: Upload
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/upload" \
-H "client_id: YOUR_CLIENT_ID" \
-H "client_secret: YOUR_CLIENT_SECRET" \
-F "file=@invoice_batch.pdf"
# {"documentId":"doc_abc123"}
# Step 2: Start extraction
curl -X POST "https://na1.fusion.foxit.com/pdf-services/api/documents/pdf-structural-extract" \
-H "client_id: YOUR_CLIENT_ID" \
-H "client_secret: YOUR_CLIENT_SECRET" \
-H "Content-Type: application/json" \
-d '{"documentId":"doc_abc123"}'
# 202 Accepted: {"taskId":"task_xyz789"}
# Step 3: Poll task status
curl "https://na1.fusion.foxit.com/pdf-services/api/tasks/task_xyz789" \
-H "client_id: YOUR_CLIENT_ID" \
-H "client_secret: YOUR_CLIENT_SECRET"
# {"taskId":"task_xyz789","status":"COMPLETED","progress":100,"resultDocumentId":"result_def456"}
# Step 4: Download the result ZIP
curl "https://na1.fusion.foxit.com/pdf-services/api/documents/result_def456/download" \
-H "client_id: YOUR_CLIENT_ID" \
-H "client_secret: YOUR_CLIENT_SECRET" \
-o extraction_result.zip The Python version with a polling loop and ZIP parsing:
import requests, json, time, zipfile
BASE_URL = "https://na1.fusion.foxit.com/pdf-services/api"
HEADERS = {"client_id": "YOUR_CLIENT_ID", "client_secret": "YOUR_CLIENT_SECRET"}
# Step 1: Upload
with open("invoice_batch.pdf", "rb") as f:
doc_id = requests.post(
f"{BASE_URL}/documents/upload", headers=HEADERS, files={"file": f}
).json()["documentId"]
# Step 2: Start extraction
task_id = requests.post(
f"{BASE_URL}/documents/pdf-structural-extract",
headers={**HEADERS, "Content-Type": "application/json"},
json={"documentId": doc_id},
).json()["taskId"]
# Step 3: Poll until COMPLETED or FAILED
while True:
task = requests.get(f"{BASE_URL}/tasks/{task_id}", headers=HEADERS).json()
if task["status"] == "COMPLETED":
result_doc_id = task["resultDocumentId"]
break
if task["status"] == "FAILED":
raise RuntimeError(f"Extraction failed: {task.get('error')}")
time.sleep(2)
# Step 4: Download the result ZIP and save it locally for inspection,
# then parse StructureInfo.json from the saved file
response = requests.get(
f"{BASE_URL}/documents/{result_doc_id}/download", headers=HEADERS
)
with open("advanced-extraction-result.zip", "wb") as f:
f.write(response.content)
with zipfile.ZipFile("advanced-extraction-result.zip") as zf:
json_name = next(n for n in zf.namelist() if n.endswith("StructureInfo.json"))
result = json.loads(zf.read(json_name))["analyzeResult"]
print(f"Schema: {result['version']['schema']}, Elements: {len(result['elements'])}")
On a clean run you should see output like Schema: 1.0.7, Elements: 9 for a small invoice batch. You’ll also find a fresh advanced-extraction-result.zip next to your script. That ZIP holds the full API response, including StructureInfo.json and any rendered image or table binaries, so you can inspect everything the engine returned and not just the parsed JSON.
First, set up and activate a Python virtual environment in your project folder. The official venv guide covers the exact commands for macOS, Linux, and Windows.
Once the virtualenv is active, the sample only needs one third-party package. Drop this into a requirements.txt next to your script and install it with pip install -r requirements.txt:
requests>=2.31.0
If you’re on macOS, use Homebrew Python (brew install python) rather than the system Python from the Xcode command-line tools. The Xcode build is linked against LibreSSL, which is enough to make a correct sample fail.
The ZIP contains a StructureInfo.json file whose top-level object wraps everything under analyzeResult. Inside that wrapper you get a version object, a pages array, a flat elements array, and an info block with analysis metadata. Each element carries its own id, type, content, region (with page and an 8-point boundingBox polygon [x1,y1,x2,y2,x3,y3,x4,y4]), and a score confidence value:
{
"analyzeResult": {
"version": {
"schema": "1.0.7",
"software": "FoxitPDFAnalyzer",
"model": "idp-analysis"
},
"pages": [
{
"pageNumber": 1,
"size": { "width": 612, "height": 792, "unit": "point" },
"state": "success"
}
],
"elements": [
{
"id": "title1",
"type": "title",
"content": {
"text": "Q3 Revenue Summary",
"style": {
"fontName": "Helvetica",
"fontSize": 24.0,
"fontWeight": 0,
"fontItalic": false
}
},
"region": {
"page": 1,
"boundingBox": [72, 47, 317, 47, 317, 80, 72, 80]
},
"score": 0.76
}
],
"info": {
"basicInfo": {
"softwareVersion": "1.6.0",
"analyzedPageCount": 1,
"elementCounts": { "title": 1 }
},
"extendedMetadata": {
"pageCount": 1,
"isEncrypted": false,
"hasAcroform": false,
"language": "en"
}
}
}
} Elements of type table, image, and form carry additional type-specific payload on top of this base shape, and any rendered image or table binary lands as a sibling file inside the ZIP referenced from the element.
HTTP errors return a standard error envelope:
{ "code": "VALIDATION_ERROR", "message": "documentId is required" } The documented error codes include VALIDATION_ERROR (400), MAX_UPLOAD_SIZE_EXCEEDED (413), DOCUMENT_NOT_FOUND (404), STORAGE_ERROR, and INTERNAL_SERVER_ERROR (500).
Password-protected PDFs that arrive with no password parameter reach the processing stage before failing. That failure surfaces in the task status poll response after status reaches FAILED, so your error handler must inspect the task response body in addition to the HTTP status codes from the initial POST calls:
{
"taskId": "task_xyz789",
"status": "FAILED",
"progress": 0,
"error": {
"code": "INTERNAL_SERVER_ERROR",
"message": "Document is password-protected"
}
} Wiring Extracted PDF Data Into Your Workflow
Pattern 1: AI/RAG pipeline. Filter the flat elements array to title, head, and paragraph types. Chunk by heading hierarchy, iterating over the array in the order the engine returned it (document reading order is preserved across columns and pages). Embed each chunk and index in Pinecone, pgvector, or your vector store of choice. Correct reading order, as provided by the extraction engine, is the prerequisite for accurate RAG retrieval on multi-column and paginated documents. When chunks split mid-thought because a layout detector merged two columns, retrieval recall drops and answer quality follows.
Pattern 2: BI reporting. Filter elements by type == "table" client-side, then convert each table’s cell structure into a pandas DataFrame:
import pandas as pd
# `result` is the `analyzeResult` object loaded from StructureInfo.json
tables = [e for e in result["elements"] if e["type"] == "table"]
for i, tbl in enumerate(tables):
# Cells live at content.body.cells[]. Each cell carries rowIndex,
# columnIndex, and a nested paragraph whose content.text holds the value.
body = tbl["content"]["body"]
grid = [["" for _ in range(body["columnCount"])] for _ in range(body["rowCount"])]
for cell in body.get("cells", []):
text = cell.get("paragraph", {}).get("content", {}).get("text", "")
grid[cell["rowIndex"]][cell["columnIndex"]] = text
df = pd.DataFrame(grid[1:], columns=grid[0]) # first row as header
print(f"Table {i}: {df.shape[0]} rows x {df.shape[1]} cols")
# df.to_gbq("finance.q3_revenue", project_id="your-project") # BigQuery
# df.to_sql("q3_revenue", engine) # Postgres / Snowflake The row and column indices from the extraction schema map directly to DataFrame positions, so you get a correctly-structured table with zero manual parsing.
Pattern 3: n8n automation. The four-step flow maps to a chain of HTTP Request nodes in n8n. The first node uploads to POST .../upload and passes documentId through the item. The second sends POST .../pdf-structural-extract and captures taskId. A Loop Over Items construct with an HTTP Request node calling GET .../tasks/{taskId} on a two-second interval checks status until COMPLETED, then routes to the download node. The final HTTP Request node calls GET .../documents/{resultDocumentId}/download, and a Code node using n8n’s binary data helpers unpacks the ZIP and parses the JSON for routing to a Salesforce, HubSpot, Postgres, or Airtable node. The polling requirement makes this a multi-node workflow, but you write zero custom glue code and gain n8n’s built-in error routing and retry handling.
PDF Extraction Tools Compared: Foxit vs. Adobe, Google, Amazon, and Azure
| Tool | Underlying Approach | Ecosystem Lock-in | Handles Scanned PDFs | Pricing Model | Setup Overhead | Status |
|---|---|---|---|---|---|---|
| Foxit Structural Extraction | Proprietary OCR + layout recognition + AI (integrated core engine) | Cloud-agnostic REST API | Yes (dedicated OCR layer) | Subscription, no per-page credits | Low (2 credential headers, 4 REST calls) | Trial (schema v1.0.7) |
| Adobe PDF Extract API | Adobe Sensei ML, reading order + renditions | Adobe Document Services | Yes | Contact sales | Medium (Adobe SDK + ecosystem) | GA |
| Google Document AI | Cloud ML + generative AI, Document Object Model | Google Cloud required | Yes | Per-page pay-as-you-go | Medium-high (GCP + IAM) | GA |
| Amazon Textract | Deep learning OCR, key-value and table extraction | AWS-native | Partial (strong on forms, weaker on complex layouts) | Per-page pay-as-you-go | Medium (AWS + IAM) | GA |
| Azure Document Intelligence | Prebuilt + custom ML models | Azure ecosystem | Yes (prebuilt models) | Per-page + model training costs | High for custom models | GA |
Google Document AI and Azure Document Intelligence win on ecosystem integration if you’re all-in on those clouds. Adobe wins on PDF structural fidelity for workflows already inside the Adobe Document Services ecosystem. Amazon Textract excels on standardized form documents where its pre-trained schema fits the input. These are real advantages, and the comparison is honest only when those contexts are acknowledged.
Foxit’s case is strongest when you need a cloud-agnostic REST API with zero ecosystem dependency, full object coverage across all twelve element types, and enterprise throughput (10 to 10,000+ PDFs/day) with SOC 2, GDPR, and HIPAA compliance built in. The Structural Extraction status is a real trade-off to factor in. The schema at v1.0.7 is callable and stable enough for pipeline integration today, but GA competitors carry a finalized contract. Pin your parser to the version field in the response and you’re insulated from schema evolution.
Your First PDF Extraction API Call, Right Now
Go to developer-api.foxit.com, create a free developer account (no credit card required), and copy your Client ID and Client Secret from the default application. Use the built-in API Playground or import the Postman collection from the Developer Portal to run the four-step sequence: upload a real document (an invoice, a multi-page contract, or a scanned form), call pdf-structural-extract with the returned documentId, poll tasks/{taskId} until COMPLETED, then download via documents/{resultDocumentId}/download.
Unzip the result, open StructureInfo.json, and check three things: analyzeResult.version.schema should report 1.0.7, analyzeResult.elements[] should contain at least one table element and one form element if your source document includes those, and the ZIP root should contain the corresponding binary files for any image-type elements. That verification confirms the full extraction pipeline is wired correctly end-to-end.
The same endpoint pattern scales to enterprise volumes. Increase upload and poll concurrency horizontally and the architecture stays identical, with no schema changes, no infrastructure modifications, and no per-page credit consumption to track.
The engineering gap between what basic extraction libraries return and what downstream systems actually consume is where document pipeline hours accumulate. Structural Extraction closes that gap at the API layer, so the complexity stays in the engine and out of your codebase. Get started at developer-api.foxit.com.
PDF Structural Extraction FAQ
What is PDF structural extraction?
PDF structural extraction is the process of identifying and classifying the semantic elements inside a PDF, such as titles, paragraphs, tables, forms, images, and annotations, rather than just pulling raw text. Foxit’s PDF Structural Extraction API returns twelve distinct element types as structured JSON, preserving spatial relationships, reading order, and table cell grids so downstream systems like RAG pipelines, BI dashboards, and CRMs can consume the data without manual parsing.
Can Foxit's API extract text from scanned PDFs?
Yes. Foxit’s PDF Structural Extraction engine includes a dedicated OCR layer that recognizes characters from image-based and scanned PDFs across 200+ languages. The OCR runs on the same internal page representation as the rendering engine, so it handles edge cases like text overlapping image regions, stamped signatures, and engineering drawing annotations that basic libraries like PyMuPDF silently drop.
How does Foxit's PDF extraction API differ from Adobe, Google Document AI, and Amazon Textract?
Foxit’s API is cloud-agnostic with no ecosystem lock-in, requiring just two credential headers and four REST calls. Adobe PDF Extract requires the Adobe Document Services ecosystem, Google Document AI requires GCP and IAM setup, and Amazon Textract requires AWS infrastructure. Foxit also uses subscription-based pricing without per-page credits, while Google, AWS, and Azure all charge per page.
What PDF elements can Foxit's Structural Extraction API identify?
The API identifies twelve element types: title, head, paragraph, table, image, headerFooter, form, hyperlink, footnote, sidebar, annotation, and formula. Each element returns with its content, an 8-point bounding box polygon, page location, and a confidence score. Tables include full cell grids with row and column indices, forms include field data, and images are extracted as separate binary files inside the result ZIP.
How do I call the Foxit PDF Structural Extraction API?
The API uses a four-step asynchronous flow: upload the PDF via POST /documents/upload to get a documentId, start extraction with POST /documents/pdf-structural-extract, poll GET /tasks/{taskId} every two seconds until status is COMPLETED, then download the result ZIP via GET /documents/{resultDocumentId}/download. Authentication uses two headers, client_id and client_secret, available from the default application in the Foxit Developer Portal.
Is the Foxit PDF Structural Extraction API ready for production use?
The endpoint is currently in Trial status with schema version v1.0.7, meaning the contract is stable but may evolve. It runs on the production base URL at developer-api.foxit.com and is built on Foxit’s core PDF engine, which powers 700 million+ users across 20+ years of deployments. For production pipelines, pin your parser to the version field in the response to insulate against future schema changes.
Automate Dynamic PDF Generation with the Foxit DocGen API: Word Templates, JSON Data, and Real API Calls

Skip the HTML-to-PDF headaches. Use Foxit’s DocGen API to turn Word templates and JSON data into clean, formatted PDFs with one API call.
If you’ve tried to generate a contract or invoice from HTML, you’ve probably burned hours on page-break-inside: avoid declarations that Chrome renders one way and a headless browser renders another. Headers and footers require separate print-media queries, and by the time you’ve got a repeating table header working correctly across pages, you’ve invested a full day of engineering into CSS that exists solely to trick a browser into behaving like a printer.
HTML documents reflow content into a viewport while PDF documents have fixed page geometry. Forcing one model into the other produces predictable failure modes: footnotes that collide with page footers, tables that split at the worst possible row, custom fonts that substitute silently, and signature blocks that drift off-page on longer documents.
There’s a larger practical cost too. For most teams, the authoritative source for enterprise document templates is already a Word file. Your legal team owns the NDA in .docx format. Finance owns the invoice in .docx format. Every structural change flows through Word because that’s where the tracked changes, formatting history, and review process live. Maintaining a parallel HTML version of each template doubles your maintenance surface from day one.
Foxit’s DocGen API eliminates that parallel entirely. You keep your templates as .docx files, embed data tags directly in Word, POST the base64-encoded template and a JSON payload to a single REST endpoint, and receive the rendered PDF (or DOCX) in the response body. You eliminate the browser rendering engine, the print-media CSS layer, and the overhead of a second template format.
How the Foxit DocGen API Works
The core model is a single synchronous POST to the GenerateDocumentBase64 endpoint at developer-api.foxit.com. Your request body carries three fields:
base64FileString: your .docx template, base64-encodeddocumentValues: a JSON object containing your merge dataoutputFormat: either"pdf"or"docx"
The API processes the template, resolves every tag against your data, and returns a JSON response containing base64FileString (the rendered document) and a message field confirming success or describing a failure. The exchange is fully synchronous, so you receive the finished document in the same HTTP response with no job ID to poll and no webhook to configure.
Authentication uses two HTTP headers: client_id and client_secret. Both come from the Foxit Developer Portal when you create an account. The free Developer plan provides 500 credits per year with no credit card required, and each GenerateDocumentBase64 call consumes exactly one credit. The Startup plan ($1,750/year) provides 3,500 credits. The Business plan ($4,500/year) covers 150,000 credits for production workloads. For context, Nutrient’s API starts at $75 for 1,000 credits, and Apryse requires a sales conversation before you can access pricing at all.
The complete call flow runs from template file to PDF on disk.
You can explore every endpoint in the live API playground at developer-api.foxit.com, and the portal includes a Postman collection you can import to run authenticated requests without writing a line of code first.
Build a Word Template with DocGen Tags
Open any .docx file in Microsoft Word and type your tags as plain text directly in the document. The DocGen API uses double-brace syntax: {{field_name}}. Tags go anywhere Word accepts text: headings, body paragraphs, table cells, headers, footers, or text boxes.
Scalar field tags resolve directly to the matching key from your documentValues JSON. A document header with {{customer_name}}, {{invoice_number}}, and {{invoice_date}} pulls those three values straight from the top-level keys of your payload.
For arrays, you wrap a single table row (the data row, not the header row) with {{TableStart:array_name}} and {{TableEnd:array_name}} markers. The wrapped row acts as a template row, and the API renders one output row per item in the JSON array. An invoice line-items table in Word looks like this:
| Description | Qty | Unit Price | Total |
|---|---|---|---|
{{TableStart:line_items}}{{description}} | {{qty}} | {{unit_price}} | {{total}}{{TableEnd:line_items}} |
Within the array row, ROW_NUMBER auto-increments with each rendered row. A SUM(ABOVE) field placed in the row directly below the {{TableEnd:line_items}} marker calculates a column total across all rendered data rows.
For nested JSON objects, use dot-notation in your tags. A shipping address block references {{shipping.street}}, {{shipping.city}}, and {{shipping.postal_code}}, mapping to properties nested inside a shipping object in your payload. The nesting can go multiple levels deep, so {{customer.address.city}} resolves against documentValues.customer.address.city.
For a working starting point, grab the downloadable invoice template from the foxit-demo-templates repo. The file is well under the 4 MB upload limit and demonstrates every pattern this article uses: scalar tags, {{TableStart:line_items}} / {{TableEnd:line_items}} with {{ROW_NUMBER}}, currency and date format switches, and subtotal / tax / total fields below the line-items table.
One sizing constraint applies while you build your own template. DocGen rejects uploads larger than 4 MB, so if you embed product photos, scanned letterhead, or full font subsets, compress the images before saving, drop embedded fonts where you can rely on system fonts, or split a large template into smaller per-section templates that you generate and merge separately.
Make Your First API Call: Generate a PDF from JSON
Run a quick pre-flight check before the first call to catch the issues that derail most clean-account run-throughs:
- Account created and
client_id/client_secretcopied from the Developer Portal API Keys section - Sample template saved locally as
invoice_template.docxin the directory you’ll run the script from - Template file size confirmed under 4 MB (
ls -lh invoice_template.docxon macOS or Linux, right-click → Properties on Windows)
With those in place, confirm your credentials work with a cURL call. The Foxit Developer Portal includes a Postman collection for this, but a quick cURL request against the API catches auth issues before any code runs:
curl -X POST "https://na1.fusion.foxit.com/document-generation/api/GenerateDocumentBase64" \
-H "client_id: YOUR_CLIENT_ID" \
-H "client_secret: YOUR_CLIENT_SECRET" \
-H "Content-Type: application/json" \
-d '{"base64FileString":"","documentValues":{},"outputFormat":"pdf"}' A 401 here means invalid credentials. A 400 with a message about the template confirms your headers are accepted and you can proceed to the full call.
Save your .docx template as invoice_template.docx in the same directory as this script, then run the complete generation:
import requests
import base64
CLIENT_ID = "your_client_id"
CLIENT_SECRET = "your_client_secret"
API_URL = "https://na1.fusion.foxit.com/document-generation/api/GenerateDocumentBase64"
# Read and encode the template
with open("invoice_template.docx", "rb") as f:
template_b64 = base64.b64encode(f.read()).decode("utf-8")
# Build the data payload
document_values = {
"customer_name": "Acme Corporation",
"invoice_number": "INV-2025-0042",
"invoice_date": "07/15/2025",
"due_date": "08/14/2025",
"line_items": [
{
"description": "API Integration Consulting",
"qty": 8,
"unit_price": 195.00,
"total": 1560.00
},
{
"description": "Document Automation Setup",
"qty": 1,
"unit_price": 750.00,
"total": 750.00
}
],
"subtotal": 2310.00,
"tax_rate": 0.08,
"tax_amount": 184.80,
"total_due": 2494.80
}
# Construct the request body
payload = {
"base64FileString": template_b64,
"documentValues": document_values,
"outputFormat": "pdf"
}
headers = {
"client_id": CLIENT_ID,
"client_secret": CLIENT_SECRET,
"Content-Type": "application/json"
}
response = requests.post(API_URL, json=payload, headers=headers)
if response.status_code == 200:
result = response.json()
pdf_bytes = base64.b64decode(result["base64FileString"])
if pdf_bytes[:5] != b"%PDF-":
raise ValueError("Response did not contain a valid PDF")
with open("invoice_output.pdf", "wb") as out:
out.write(pdf_bytes)
print("PDF written to invoice_output.pdf")
else:
print(f"Error {response.status_code}: {response.json().get('message')}") The success response is a JSON object with three keys: base64FileString (the rendered PDF, base64-encoded), fileExtension ("pdf"), and message ("PDF Document Generated Successfully"). Decoding and writing the bytes to disk gives you a complete, formatted PDF with every tag replaced by its corresponding data value. If you omit a key from documentValues, the API renders the corresponding tag as an empty string, producing a blank field in the output.
Advanced Data Scenarios: Arrays, Nested Objects, and Built-In Functions
The two-row invoice above works, but most production documents have more complex data shapes. Three patterns cover the majority of real-world cases.
For multi-row tables, the line_items array in the Python snippet above already shows the basic structure. To generate five rows, pass five objects in the array. The Word template row tagged with {{TableStart:line_items}} and {{TableEnd:line_items}} repeats exactly once per array item:
{
"line_items": [
{
"description": "UX Design Review",
"qty": 4,
"unit_price": 150.0,
"total": 600.0
},
{
"description": "Backend API Development",
"qty": 12,
"unit_price": 185.0,
"total": 2220.0
},
{
"description": "Database Schema Migration",
"qty": 3,
"unit_price": 200.0,
"total": 600.0
},
{
"description": "QA Testing",
"qty": 6,
"unit_price": 95.0,
"total": 570.0
},
{
"description": "Deployment and Documentation",
"qty": 2,
"unit_price": 175.0,
"total": 350.0
}
]
} The API generates exactly five table rows. Swap in 50 items and you get 50 rows, with page breaks handled by Word’s native pagination logic.
For nested objects, the DocGen API resolves dot-notation paths against the full depth of your JSON structure. A shipping confirmation template referencing {{customer.address.city}} works against this payload without any flattening on your end:
{
"customer": {
"name": "Sarah Chen",
"email": "[email protected]",
"address": {
"street": "742 Evergreen Terrace",
"city": "Portland",
"state": "OR",
"postal_code": "97201"
}
}
} In the Word template, {{customer.name}}, {{customer.address.city}}, and {{customer.address.postal_code}} each resolve to the correct nested value. You can reference the same nested object from multiple locations in the template, and the API populates each instance independently.
For numeric and date formatting, the DocGen API respects Word’s native field switch syntax. Adding \# Currency to a tag formats a numeric value as a currency string, so {{unit_price \# Currency}} renders 195.00 as \$195.00. Date fields accept \@ "MM/dd/yyyy" to control output format, so {{invoice_date \@ "MM/dd/yyyy"}} formats an ISO date string to 07/15/2025. To auto-calculate a column total, place a SUM(ABOVE) field in the Word table row immediately below {{TableEnd:line_items}} and the API evaluates it against the rendered data rows.
Error Handling and Production Readiness
The DocGen API returns a focused set of HTTP status codes. A 200 confirms successful generation. A 401 means your client_id or client_secret headers are invalid, and the fix is to re-copy the credentials from the Developer Portal. A 400 covers three cases. The first is a malformed request body, for example a missing base64FileString or outputFormat. The second is structural issues with the template itself, such as a {{TableStart}} marker placed outside its table row. The third is an oversize template; DocGen rejects .docx uploads larger than 4 MB, and the fix is to compress embedded images, drop embedded fonts, or split the template before re-encoding. The message field in every non-200 response body gives you the specific reason, so log it rather than discarding the response object.
A production wrapper handles all three cases and adds exponential backoff for transient server errors:
import requests
import base64
import time
def generate_document(client_id, client_secret, template_path,
document_values, output_format="pdf"):
API_URL = "https://na1.fusion.foxit.com/document-generation/api/GenerateDocumentBase64"
with open(template_path, "rb") as f:
template_b64 = base64.b64encode(f.read()).decode("utf-8")
payload = {
"base64FileString": template_b64,
"documentValues": document_values,
"outputFormat": output_format
}
headers = {
"client_id": client_id,
"client_secret": client_secret,
"Content-Type": "application/json"
}
max_retries = 3
for attempt in range(max_retries):
try:
response = requests.post(API_URL, json=payload,
headers=headers, timeout=30)
if response.status_code == 200:
return base64.b64decode(response.json()["base64FileString"])
if response.status_code == 401:
raise ValueError("Authentication failed: re-check client_id and client_secret")
if response.status_code == 400:
msg = response.json().get("message", "Bad request")
raise ValueError(f"Request error: {msg}")
if response.status_code >= 500:
if attempt < max_retries - 1:
wait = 2 ** attempt
print(f"Server error ({response.status_code}), retrying in {wait}s...")
time.sleep(wait)
continue
raise RuntimeError(f"Server error after {max_retries} attempts")
except requests.exceptions.Timeout:
if attempt < max_retries - 1:
time.sleep(2 ** attempt)
continue
raise
raise RuntimeError("Max retries exceeded") The wrapper raises immediately on 4xx responses because retrying a credential error or a malformed request produces the same result. Exponential backoff applies only to 5xx responses and timeouts, where the issue is transient.
Once generate_document() returns raw PDF bytes, routing them downstream takes three lines:
import boto3
s3 = boto3.client("s3")
pdf_bytes = generate_document(CLIENT_ID, CLIENT_SECRET, "invoice_template.docx", document_values)
s3.put_object(Bucket="my-documents-bucket", Key="invoices/INV-2025-0042.pdf", Body=pdf_bytes) To attach the output to an email, pass pdf_bytes directly as the smtplib attachment payload. To collect a signature on the generated document, base64-encode the bytes and POST them to Foxit’s eSign API with the signer’s email address in the request body. The full eSign API reference is at docs.developer-api.foxit.com.
Common Mistakes
A short list of the issues that account for almost every failed first run.
- Smart-quote autocorrect on braces. Word’s AutoCorrect can convert the second
{of{{into a curly-quote glyph, which breaks tag parsing silently. Disable “Straight quotes with smart quotes” under AutoCorrect Options, or paste tags as plain text. - Token case sensitivity.
{{Customer_Name}}and{{customer_name}}are different keys. Match the casing in your JSON exactly. TableStartandTableEndmust sit in the same Word table row. Splitting them across two rows, or placing either marker outside the table, leaves the loop unrendered with no error.- Template over 4 MB. The API rejects oversize uploads with a 400. Compress embedded images, drop embedded fonts where system fonts will do, or split the template into smaller pieces.
- Missing payload key. The API renders an unmatched tag as an empty string rather than failing, so a 200 response does not guarantee every field is populated. Spot-check the rendered PDF as part of any pipeline test.
- Auth header typos. Headers are
client_idandclient_secretin snake_case.Client-Id,ClientId, orX-Client-Idall return 401.
Run the Full Invoice Example End-to-End Right Now
Create a free account directly at account.foxit.com/site/sign-up. This skips the pricing-page redirect you hit from the marketing site and drops you straight into the account form.
- Open account.foxit.com/site/sign-up and complete the form (no credit card required).
- After verification, sign in to the Developer Portal and the Developer plan (500 credits per year) is active by default.
- Open the API Keys section and copy your
client_idandclient_secret.
With credentials in hand, run the example end-to-end:
- Download
invoice_full.docxfrom the foxit-demo-templates repo and save it locally asinvoice_template.docxin your working directory. The file is well under the 4 MB upload limit and exercises every tag pattern this article covers. - Paste your credentials into the
CLIENT_IDandCLIENT_SECRETvariables in the Python script from the previous section. - Edit the
document_valuesdictionary with your own customer name, invoice number, and line items. - Run the script and open
invoice_output.pdf.
The free Developer plan’s 500 annual credits cover this tutorial dozens of times over before you spend anything. The full API reference at docs.developer-api.foxit.com covers every endpoint parameter, the complete tag specification, all supported output formats, and the full GenerateDocumentBase64 request and response schema.
Get started with a free account (no credit card required) and generate your first dynamic PDF in under 10 minutes.
Document Workflow Automation: An Architectural Guide to Building API-Driven Document Pipelines

Automate document workflows with APIs. Learn how to scale PDF generation, eSign, and processing pipelines using modern architecture.
A PDF generation script that breaks on special characters. A cron job that retries failed document conversions by rerunning the entire job. An eSign flow tracked in a shared spreadsheet where “sent” means someone sent an email. These aren’t hypothetical failure modes; they’re the actual engineering artifacts that accumulate when document workflows grow faster than the architecture beneath them.
The scale problem compounds quickly. A team processing 200 contracts a month can survive on scripts and email hand-offs. At 2,000 contracts, those same workflows are the bottleneck. At 20,000, engineers are maintaining hacks that should have been replaced two years ago: retry logic bolted onto cron jobs, signing flows with no audit trail, and PDF generation that silently drops content when a CRM field contains a Unicode character.
The global intelligent document processing market was valued at $2.3B in 2024 and is projected to reach $12.35B by 2030 at a 33.1% CAGR, not because AI is newly fashionable, but because manual document handling is a measurable operational ceiling. The organizations crossing that ceiling aren’t doing it by adopting better tools in isolation. They’re adopting an architectural model.
The problem isn’t a lack of API options for document generation, conversion, or signing. The problem is the absence of a framework for assembling those operations into a pipeline that’s resilient, auditable, and testable. This guide gives you that framework, then grounds it in working Python examples against a real REST API suite.
Anatomy of a Document Automation Pipeline: The Five Stages
Before you write a single API call, you need a model for what you’re building. Every document workflow automation pipeline, regardless of domain, decomposes into five discrete stages.
Stage 1 is intake: you receive or capture the source data that will drive the document. This might be a webhook payload from your CRM when a deal closes, a form submission, or a batch export from an ERP system. The manual failure mode here is no schema validation, no deduplication, and no observable queue depth. Documents arrive out of order, get processed twice, or disappear without trace.
Stage 2 is generation: you render a document from a template and the structured data from stage 1. Common outputs include contracts, invoices, compliance reports, and onboarding kits. The failure mode is template version drift (production runs a different template version than staging), no validation of input data against the template’s expected schema, and no idempotent retry path if the generation call fails partway through.
Stage 3 is processing: you transform, extract from, or optimize the generated document. This covers format conversion (DOCX to PDF), content extraction for downstream indexing, compression, and linearization for fast web delivery. The failure mode is processing steps chained with no error isolation, so a failed compression step blocks the entire document from reaching signing.
Stage 4 is signing: you route the document for signature, track signer status, and capture consent with a full audit trail. The failure mode is manual polling for signer status, no webhook-driven callbacks, and no programmatic access to the audit log when a compliance review is triggered.
Stage 5 is archival and distribution: you store the signed document with a retention policy and push it to downstream systems, your DMS, CRM, or data warehouse. The failure mode is no content-addressed versioning, no record of which document version was signed, and no delivery confirmation to downstream consumers.
Idempotency is a first-class requirement at every stage. Each operation should be safely retryable: the same inputs produce the same output, and a retried call doesn’t create a duplicate document, signing request, or archive record. You implement idempotency in your orchestration layer by generating a unique key per document job and checking it before re-processing. This is a design responsibility. The API doesn’t handle it for you automatically.
The data flow through a well-designed document automation pipeline looks like this:

One constraint to know upfront: the three APIs in this stack don’t share a document ID namespace. Each stage boundary requires a file handoff. DocGen returns the rendered document as base64 in the response body. You decode it and either save it to disk or upload it directly to PDF Services. PDF Services returns a resultDocumentId that you download as a file, then re-upload to eSign, which runs on a different host with different authentication. The handoff pattern is a feature, not a limitation. It makes each stage independently testable and replayable.
Architectural Decision Framework: Four Axes Before You Write Code
Four decisions determine whether your document pipeline scales cleanly or becomes the thing your team rewrites in 18 months.
Axis 1: REST API vs. SDK
Use REST APIs for cloud-native, horizontally scalable pipelines where document operations are stateless HTTP calls. Use an SDK for on-premise deployments, air-gapped environments, or latency-sensitive processing where network round-trips are a constraint. Foxit offers both: REST APIs for cloud-native pipelines and PDF SDKs for on-premise or air-gapped deployments, so the axis is a real choice, not a theoretical one. If your document pipeline runs inside a regulated environment where data can’t leave the network perimeter, the SDK is the correct answer regardless of how convenient the REST API is.
Axis 2: Synchronous vs. Asynchronous Processing
This is the most consequential call you’ll make, and it varies by stage within a single pipeline.
| Factor | Synchronous | Asynchronous |
|---|---|---|
| Document size | Under ~10 pages | Large or variable-length |
| SLA requirement | Sub-second response | Variable completion time acceptable |
| Typical use case | Real-time contract preview | Batch invoice processing |
| Error handling | Inline exception handling | Dead-letter queue, retry on callback |
| Foxit API example | DocGen (returns document in response body) | PDF Services (returns taskId, poll for result); eSign (webhook callback on folder execution) |
The Foxit suite itself illustrates this split cleanly. DocGen is synchronous: POST your template and data payload, get the rendered document back immediately in the response body. No taskId, no polling. PDF Services is asynchronous: a conversion call returns a taskId, and you poll a status endpoint until the result is ready. eSign is asynchronous via webhooks: creating a folder returns immediately, and the API delivers a callback to your registered endpoint when the folder is executed (all signers complete). Design your pipeline around this reality rather than assuming a uniform execution model across all three APIs.
Axis 3: Linear Pipeline vs. Event-Driven Architecture
A linear pipeline (where stage A blocks until complete before stage B starts) works for simple three-stage flows with predictable volume and acceptable end-to-end latency. An event-driven pipeline, where each stage emits a completion event consumed by the next stage, is the correct choice when you need error isolation (a failed stage 3 doesn’t block stage 2 outputs from being replayed), partial replay (reprocess from stage 2 without regenerating the document), or parallel processing branches (send the same document to multiple downstream consumers simultaneously).
For pipelines that start as linear but need to scale, n8n is a practical bridge. You can call Foxit’s REST APIs from n8n workflows via HTTP Request nodes, which lets you wire pipeline stages without writing custom glue code while you validate the workflow logic before committing to a fully coded implementation.
Axis 4: Error Handling Strategy for Document Pipelines
Three components belong in your initial design, not bolted on afterward.
The first is idempotency keys. Generate a unique key per document job (a UUID tied to the source record ID and timestamp works well) and check it before re-processing. If a worker crashes mid-job and the job re-queues, the idempotency key prevents duplicate processing.
The second is dead-letter handling. Define what happens to a document that has failed three consecutive processing attempts. It should route to a dead-letter queue with the failure reason and enough context to replay it manually or trigger an alert.
The third is a circuit breaker. If PDF Services returns 5xx responses on five consecutive calls within 30 seconds, stop sending requests and return a fast failure to the calling system. This prevents a degraded upstream API from exhausting your worker pool and cascading failures downstream. The circuit breaker pattern maps cleanly onto any stateless HTTP integration.
Building the Pipeline: Foxit APIs in Practice
We’ll use Foxit’s PDF Services, DocGen, and eSign APIs for the examples below. The patterns translate to any REST-based document API, but these are the endpoints we’ll call.
Document Generation with the DocGen API
DocGen takes a DOCX template (encoded as base64) and a JSON data payload, and returns the rendered document immediately in the response body. There’s no templateId concept; you send the template inline with every request. This means you own template versioning. Keep your templates in version control and pin the version used for each job to your event log.
One practical cap to design around: the DocGen endpoint rejects .docx uploads larger than 4 MB once base64-encoded. Compress embedded images through Word’s Picture Format settings, drop embedded fonts and OLE objects, and split very large templates into multiple files before the request leaves your service.
The request uses client_id and client_secret as HTTP headers against na1.fusion.foxit.com.
# Illustrative example - not production code
import base64
import requests
import json
def generate_contract(template_path: str, data: dict) -> bytes:
with open(template_path, "rb") as f:
template_b64 = base64.b64encode(f.read()).decode("utf-8")
payload = {
"outputFormat": "pdf",
"documentValues": data,
"base64FileString": template_b64
}
response = requests.post(
"https://na1.fusion.foxit.com/document-generation/api/GenerateDocumentBase64",
headers={
"client_id": "YOUR_CLIENT_ID",
"client_secret": "YOUR_CLIENT_SECRET",
"Content-Type": "application/json"
},
json=payload
)
response.raise_for_status()
result = response.json()
return base64.b64decode(result["base64FileString"])
# Data pulled from your CRM or ERP; validate against your template schema before calling
contract_data = {
"client_name": "Acme Corp",
"contract_value": "48000",
"effective_date": "2025-09-01",
"payment_terms": "Net 30"
}
pdf_bytes = generate_contract("templates/msa_v3.docx", contract_data) Validate your data payload against the template’s expected field schema before the API call. DocGen doesn’t catch type errors or missing fields with a clean error response. You get a malformed document instead. A Pydantic model or JSON Schema validation step before the POST saves significant debugging time.
PDF Processing with the PDF Services API
The most common PDF Services operation is conversion. The DOCX-to-PDF call is also the simplest entry point for teams new to the API. PDF Services uses a two-step pattern: upload the source file first to get a documentId, then call the operation endpoint with that ID. Because operations are asynchronous, the call returns a taskId that you poll until the result is available.
# Illustrative example - not production code
import time
import requests
PDF_SERVICES_HOST = "https://na1.fusion.foxit.com"
HEADERS = {
"client_id": "YOUR_CLIENT_ID",
"client_secret": "YOUR_CLIENT_SECRET"
}
def upload_document(file_bytes: bytes, filename: str) -> str:
response = requests.post(
f"{PDF_SERVICES_HOST}/pdf-services/api/documents/upload",
headers=HEADERS,
files={"file": (filename, file_bytes, "application/octet-stream")}
)
response.raise_for_status()
return response.json()["documentId"]
def poll_task(task_id: str) -> str:
while True:
status_resp = requests.get(
f"{PDF_SERVICES_HOST}/pdf-services/api/tasks/{task_id}",
headers=HEADERS
)
status_resp.raise_for_status()
status_data = status_resp.json()
if status_data["status"] == "COMPLETED":
return status_data["resultDocumentId"]
elif status_data["status"] == "FAILED":
raise RuntimeError(f"Task failed: {status_data}")
time.sleep(2)
def download_document(document_id: str) -> bytes:
response = requests.get(
f"{PDF_SERVICES_HOST}/pdf-services/api/documents/{document_id}/download",
headers=HEADERS
)
response.raise_for_status()
return response.content
def convert_docx_to_pdf(docx_bytes: bytes) -> bytes:
doc_id = upload_document(docx_bytes, "document.docx")
response = requests.post(
f"{PDF_SERVICES_HOST}/pdf-services/api/documents/create/pdf-from-word",
headers={**HEADERS, "Content-Type": "application/json"},
json={"documentId": doc_id}
)
response.raise_for_status()
result_doc_id = poll_task(response.json()["taskId"])
return download_document(result_doc_id)
def extract_text(pdf_bytes: bytes) -> str:
doc_id = upload_document(pdf_bytes, "document.pdf")
response = requests.post(
f"{PDF_SERVICES_HOST}/pdf-services/api/documents/modify/pdf-extract",
headers={**HEADERS, "Content-Type": "application/json"},
json={"documentId": doc_id, "extractType": "TEXT"}
)
response.raise_for_status()
result_doc_id = poll_task(response.json()["taskId"])
return download_document(result_doc_id).decode("utf-8") The pdf-extract endpoint pulls text from the PDF (pass extractType as TEXT, IMAGE, or PAGE depending on what you need). Both conversion and extraction follow the same upload, execute, poll, download cycle. Feed the text output to a downstream search index so the document is queryable immediately after processing.
Signature Orchestration with the eSign API
The eSign API uses OAuth2, not header-based authentication. Your first call exchanges client_id and client_secret for a Bearer token on a separate host (na1.foxitesign.foxit.com).
# Illustrative example - not production code
import json
import requests
from flask import Flask, request as flask_request
ESIGN_HOST = "https://na1.foxitesign.foxit.com"
def get_esign_token(client_id: str, client_secret: str) -> str:
response = requests.post(
f"{ESIGN_HOST}/api/oauth2/access_token",
data={
"grant_type": "client_credentials",
"client_id": client_id,
"client_secret": client_secret
}
)
response.raise_for_status()
return response.json()["access_token"]
def create_signing_folder(token: str, pdf_bytes: bytes, signers: list) -> str:
folder_payload = {
"folderName": "MSA - Acme Corp",
"parties": [
{
"firstName": s["first_name"],
"lastName": s["last_name"],
"emailId": s["email"],
"permission": "FILL_FIELDS_AND_SIGN",
"sequence": s["sequence"]
}
for s in signers
]
}
response = requests.post(
f"{ESIGN_HOST}/api/folders/createfolder",
headers={"Authorization": f"Bearer {token}"},
files={
"file": ("contract.pdf", pdf_bytes, "application/pdf"),
"data": (None, json.dumps(folder_payload), "application/json")
}
)
response.raise_for_status()
return response.json()["folderId"]
# Webhook handler receives the folder-executed event
app = Flask(__name__)
@app.route("/webhooks/esign", methods=["POST"])
def esign_webhook():
event = flask_request.json
if event.get("event_type") == "folder_executed":
folder_id = event["folder_id"]
signed_doc_url = event["documents"][0]["download_url"]
archive_signed_document(folder_id, signed_doc_url)
return "", 200 Register your webhook endpoint in the eSign developer portal settings. When a folder is executed (all signers complete), the API POSTs the event payload to your endpoint. Extract the signed document URL from the callback and pass it to your archival stage. The eSign API also exposes a folder activity history endpoint that returns a complete audit trail: signer identity, timestamp, IP address, and authentication method for every interaction with the folder.
Chaining the Pipeline Stages with Idempotency
The file handoff between stages is explicit by design. Here’s a minimal orchestration wrapper that chains all three stages and demonstrates the idempotency pattern:
# Illustrative example - not production code
import uuid
def run_document_pipeline(job_id: str, template_path: str, data: dict, signers: list):
idempotency_key = f"{job_id}:{uuid.uuid4()}"
if is_already_processed(idempotency_key):
return # Safe to retry
# Stage 2: Generate (DocGen returns PDF bytes synchronously)
pdf_bytes = generate_contract(template_path, data)
log_pipeline_event(job_id, "generated", hash_document(pdf_bytes))
# Stage 3: Process (extract text for indexing; convert if needed)
extracted = extract_text(pdf_bytes)
index_document(job_id, extracted)
log_pipeline_event(job_id, "processed", hash_document(pdf_bytes))
# Stage 4: Sign (eSign returns folder ID; completion arrives via webhook)
token = get_esign_token("YOUR_CLIENT_ID", "YOUR_CLIENT_SECRET")
folder_id = create_signing_folder(token, pdf_bytes, signers)
log_pipeline_event(job_id, "sent_for_signature", folder_id)
mark_processed(idempotency_key) For async pipelines handling thousands of documents per hour, replace direct function calls with queue messages. Each stage worker pulls a job from Redis or Amazon SQS, executes the API call, ACKs on success, and publishes a completion event to the next stage’s queue. If a worker crashes mid-job, the unACKed message re-queues and the idempotency key prevents re-processing a document that has already been completed.
Auditability and Compliance by Design
GDPR, HIPAA, and SOC 2 Type II each impose specific requirements around document lifecycle traceability. Retrofitting an audit layer onto a pipeline that wasn’t designed for it takes far more work than building it in from the start.
The event sourcing pattern fits document pipelines directly. Maintain an append-only log of every document event: created, converted, sent_for_signature, signed, archived. Use a stable document_id as the primary key. This log makes replay straightforward: if signing fails, you can replay from the processing output without regenerating the document from scratch. Each event record should include the stage name, timestamp, operator identity, and a SHA-256 hash of the document bytes at that stage.
The SHA-256 hash at each stage isn’t overhead; it’s your tamper detection mechanism. If the hash of the document presented for signing doesn’t match the hash recorded at generation, you have an integrity problem that’s immediately visible. This satisfies document integrity requirements in regulated industries without any additional tooling.
The Foxit eSign API’s built-in audit trail captures signer identity, timestamp, IP address, and authentication method for every folder interaction. Query the folder activity history endpoint to retrieve this data and persist it in your own audit store alongside your pipeline event log. Storing it in your own system, rather than relying solely on the eSign provider’s records, gives you a complete, portable audit trail that survives a provider migration.
Scaling Document Workflow Automation Without Rebuilding It
Batch Ingestion
Place incoming document jobs on a queue (Redis list or SQS FIFO queue) and run a pool of stateless worker processes. Each worker pulls a job, executes the API call with an idempotency key, and ACKs on success. Dead-letter routing handles permanently failed documents.
This pattern processes thousands of documents per hour without hammering the API or requiring coordination between workers. Because each REST API call is stateless, workers scale horizontally without any shared state. You add capacity by adding workers, not by redesigning the pipeline.
Credit Quota and Backoff
Foxit’s pricing model is credit-based: API calls consume credits, and calls pause when credits are exhausted until renewal or upgrade. Implement exponential backoff with jitter on 5xx responses as a general practice for any REST API integration.
# Illustrative example - not production code
import time
import random
import requests
def api_call_with_retry(url, headers, payload, max_retries=4):
for attempt in range(max_retries):
response = requests.post(url, headers=headers, json=payload)
if response.status_code < 500:
return response
wait = (2 ** attempt) + random.uniform(0, 1)
time.sleep(wait)
response.raise_for_status() Log quota exhaustion as a separate metric category. Consistent credit exhaustion is a signal to upgrade your plan. It shouldn’t require digging through application logs to detect.
Observability
Instrument each pipeline stage with three metrics: processing latency (time from job enqueue to stage completion), error rate per stage, and document volume per time window. Use structured JSON logging so stage failures are queryable without parsing free-text log lines. Tools like OpenTelemetry make it straightforward to emit these metrics in a vendor-neutral format.
A document that enters the pipeline and never exits is a data integrity problem. Track in-flight documents explicitly: when a job enters signing, record it. When the eSign webhook fires, close the record. Any job that’s been in stage 4 for longer than your expected SLA without a webhook callback warrants an alert, not just a log entry.
Ship Your First Document Pipeline Stage Today
The gap between a collection of one-off scripts and a production document pipeline isn’t as wide as it looks. It starts with one stage, not five.
Create a free account directly at account.foxit.com/site/sign-up (no credit card required; the Developer plan ships with 500 credits per year). The direct URL skips the pricing-page redirect you would otherwise hit from the developer portal, so you finish on the account form and then land in the API Keys section where credentials live. From there, make your first conversion call: POST a DOCX file from your own system to the PDF Services conversion endpoint using the Python example above and confirm you get a valid PDF back. That single round-trip validates your auth, your network path, and the basic integration pattern before you write any orchestration logic.
Once that’s working, pick one document type in your system that’s currently generated or processed manually and map it to the five-stage model from the second section of this article. Find the highest-friction bottleneck stage and start there, not at stage 1. If generation is the pain point, use the Developer Playground in the developer portal to test DocGen templates against real data payloads before writing a single line of integration code. If signing is the bottleneck, wire up the eSign folder creation and a webhook handler to close the loop.
The patterns in this guide (idempotency keys, event-sourced audit logs, async stage handoffs, circuit breakers) apply to any document API stack. A unified REST API suite covering generation, processing, and signing from a single provider cuts the number of authentication models to manage, reduces integration surface area, and gives you a consistent debugging path when something fails across stages. That’s the practical payoff of treating document workflow automation as a first-class architectural concern rather than a collection of scripts that should have been replaced two years ago.
Start building your first pipeline stage today.
Frequently Asked Questions
What is document workflow automation?
Document workflow automation replaces manual, script-driven document operations (generation, conversion, signing, and archival) with a structured API-driven pipeline. Each stage is independently testable, retryable via idempotency keys, and observable through structured event logs. At scale (thousands of documents per hour), automation eliminates the bottlenecks created by cron jobs, shared spreadsheets, and one-off scripts.
When should I use a synchronous vs. asynchronous document API?
Use synchronous APIs when you need sub-second responses for small documents, for example, real-time contract previews under approximately 10 pages. Use asynchronous APIs (polling or webhook-driven) for large or variable-length documents, batch invoice processing, or any workflow where variable completion time is acceptable. Many document API suites, including Foxit’s, mix both models across different endpoints, so design each pipeline stage around the actual execution model of the specific API call it makes.
How do I make a document pipeline idempotent?
Generate a unique key per document job (a UUID tied to the source record ID and timestamp works well) and check whether that key has already been processed before executing any stage. Store processed keys in a fast key-value store (Redis is a common choice). On retry, the idempotency check returns early without duplicating the document, signing request, or archive record. This is an orchestration-layer responsibility; the document API itself doesn’t provide it automatically.
What compliance requirements apply to document pipelines?
GDPR, HIPAA, and SOC 2 Type II each require document lifecycle traceability. Implement an append-only event log keyed by a stable document_id, capturing stage name, timestamp, operator identity, and a SHA-256 hash of the document at each stage. For eSign specifically, store the provider’s audit trail (signer identity, IP address, authentication method, timestamp) in your own system so the record is portable across provider migrations.