How to Split a Multi-Document PDF Using JavaScript and Google Cloud Document AI

How to Split a Multi-Document PDF Using JavaScript and Google Cloud Document AI

Software Development
Oct 11, 2024
6-7 min

Share blog

Introduction

In this tutorial, I will guide you through a process of splitting a PDF that contains multiple documents using JavaScript, Google Cloud’s Document AI, and the pdf-lib library. This feature is useful when you have a PDF with several documents, each identified by page numbers (e.g., “Page 1 of 2” for the first document, “Page 1 of 3” for the second document, etc.). Document AI will help extract page number data, and then we’ll split the PDF accordingly.

Step 1: Understanding the Problem

Consider a PDF with multiple documents, each identified by page numbers:

  • The first document has 2 pages, labeled “Page 1 of 2”, “Page 2 of 2”.

The second document has 3 pages, labeled “Page 1 of 3”, “Page 2 of 3”, “Page 3 of 3”. We’ll use OCR (Optical Character Recognition) to extract these page numbers and split the PDF into separate files for each document.

Image

Step 2: Setting Up Google Cloud Document AI

To OCR the page numbers, we will use Google Cloud Document AI’s Custom Extractor.

1. Create a Google Cloud Account if you don’t have one.

2. Set up Document AI by searching for it in the GCP Console

Image

3. Create a Custom Processor by selecting the Custom Extractor model.

Image

4. Select Custom extractor as our processor.

Image

5. Upload Training Documents: Upload sample PDFs to train our processor .

Image

6. Create Labels: Annotate the page numbers and total page count fields, creating two labels: page_no and page_total. For optimal accuracy, label at least 100 pages across 20 documents.

Image
Image

7. Train and Deploy the model.

Step 3: Extracting Page Numbers from the PDF Using Document AI

Once the processor is trained and deployed, you can extract labeled data like page numbers and total pages from the PDF. Here’s how we do it in JavaScript:

typescript
1const name = `projects/${projectId}/locations/${location}/processors/${processorId}`;
2const buffer = await getTheArrayBufferFromPdfUrl(s3Url);
3const encodedImage = Buffer.from(buffer).toString('base64');
4
5const request = {
6 name,
7 rawDocument: {
8 content: encodedImage,
9 mimeType: 'application/pdf',
10 },
11};
12
13const [result] = await client.processDocument(request);
14const { document } = result;
15const { entities } = document;
16const pages = formatData(entities);
17const pagesToSplit = getPdfPagesToSplit(pages);

This function organizes the extracted data into a structured array containing each page’s number and total page count.

Step 4: Identifying Document Boundaries

We then determine the starting and ending pages for each document inside the PDF:

javascript
1getPdfPagesToSplit = (pages) => {
2 const pdfPages = [];
3 let count = 0;
4 let skipCount = 0;
5
6 for (const page of pages) {
7 count++;
8 if (skipCount) {
9 skipCount--;
10 continue;
11 }
12
13 if (page.page_total == 1) {
14 pdfPages.push({ number: +page.number + 1, start: count, end: count });
15 } else if (page.page_total > 1) {
16 skipCount = page.page_total - 1;
17 pdfPages.push({ number: +page.number + 1, start: count, end: count + +page.page_total - 1 });
18 }
19 }
20
21 return pdfPages;
22};
23
24 },
25};
26
27const [result] = await client.processDocument(request);
28const {document} = result;
29const {entities} = document;
30const pages = formatData(entities);
31const pagesToSplit = getPdfPagesToSplit(pages);

Step 5: Splitting the PDF Using pdf-lib

Once we have the start and end pages, we can split the PDF using pdf-lib:

javascript
1extractPdfPage = async (arrayBuff, pageToSplit) => {
2 const pdfSrcDoc = await PDFDocument.load(arrayBuff);
3 const pdfNewDoc = await PDFDocument.create();
4 const pages = await pdfNewDoc.copyPages(pdfSrcDoc, range(pageToSplit.start, pageToSplit.end));
5 pages.forEach(page => pdfNewDoc.addPage(page));
6
7 const newPdf = await pdfNewDoc.save();
8 return newPdf;
9};

Here, pdf-lib copies and saves the pages of each document as a new PDF.

Step 6: Upload or Download the Split PDFs

Now, we can take the split PDFs from SplittedPdfs and either upload them to a cloud service or download them to the user’s machine:

javascript
1const SplittedPdfs = [];
2for (const pageToSplit of pagesToSplit) {
3 const splittedPdf = await extractPdfPage(imageFile, pageToSplit);
4 SplittedPdfs.push(splittedPdf);
5}
6// Now you can use SplittedPdfs as per your needs.

Conclusion

This tutorial demonstrates how to split a multi-document PDF using JavaScript, Document AI, and pdf-lib. We covered setting up Document AI, extracting page numbers, and splitting the PDF based on those page numbers. With these steps, you can easily implement this feature in your own applications.

Blogs

Discover the latest insights and trends in technology with the Omax Tech Blog.

View All Blogs
Omax | Blog | The Right Way to Migrate from MySQL to AWS Aurora DSQL
7-8 min
August 25, 2026

The Right Way to Migrate from MySQL to AWS Aurora DSQL

Migrating a production database is one of the highest-risk changes you can make to an application. Moving from MySQL to AWS Aurora DSQL raises the stakes further...

Read More
Omax | Blog | From Memory Nightmare to Serverless: Bundling Files into a ZIP with AWS Lambda
8-10 min
August 25, 2026

From Memory Nightmare to Serverless: Bundling Files into a ZIP with AWS Lambda

A straightforward 'download all these files as one ZIP' request that worked perfectly on my laptop and... fell over the first day it met real production load. Here's the debugging story, the scaling options I ruled out, and why AWS Lambda was the right answer.

Read More
Omax | Blog | Clean Code vs. Overengineering: Where Should Developers Draw the Line?
10-12 min
August 21, 2026

Clean Code vs. Overengineering: Where Should Developers Draw the Line?

Clean code reduces unnecessary complexity; overengineering invents it. A practical guide to using context, evidence, and the cost of change to know when to stop adding abstractions...

Read More
Omax | Blog | Kafka vs RabbitMQ vs AWS EventBridge: Choosing the Right Architecture Based on Business Requirements
10-12 min
August 21, 2026

Kafka vs RabbitMQ vs AWS EventBridge: Choosing the Right Architecture Based on Business Requirements

Compare Kafka, RabbitMQ, and AWS EventBridge based on scalability, routing, event streaming, replay, infrastructure, and business requirements to choose the right architecture...

Read More
Omax | Blog | AI Integrations for QA Engineers
15-20 min
August 20, 2026

AI Integrations for QA Engineers

Learn how QA engineers can connect AI with Jira, GitHub, Slack, Notion and other tools to improve testing, bug tracking, reporting and QA productivity...

Read More
Omax | Blog | The Ultimate Guide to Amazon SES Setup with GoDaddy DNS
8-10 min
August 18, 2026

The Ultimate Guide to Amazon SES Setup with GoDaddy DNS

Learn how to set up Amazon SES with GoDaddy DNS. Complete step-by-step guide covering Easy DKIM, SPF, DMARC, custom MAIL FROM, and exiting the SES Sandbox...

Read More
Omax | Blog | AWS DevOps Agent Setup Guide with EC2
8-10 min
August 17, 2026

AWS DevOps Agent Setup Guide with EC2

Learn how to set up AWS DevOps Agent with EC2, CloudWatch, IAM, and Agent Spaces for AI-assisted monitoring, incident investigation, and root-cause analysis...

Read More
Omax | Blog | Multi-Tenancy Patterns in DynamoDB: Silo, Pool, and Bridge Models
6-10 min
August 13, 2026

Multi-Tenancy Patterns in DynamoDB: Silo, Pool, and Bridge Models

If you've already made the jump from a relational database to DynamoDB see our guide on moving relational data from SQL to DynamoDB...

Read More
Omax | Blog | We stopped leaving the IDE to design. Here’s our Cursor → Figma flow
8-10 min
August 10, 2026

We stopped leaving the IDE to design. Here’s our Cursor → Figma flow

Cursor drafts fast, catches gaps early, and still clips fields and breaks layouts. Here's the real pros-and-cons breakdown of our workflow...

Read More

Ready to Work With Us?

Most engagements start with a 20-minute conversation. No pitch, no pressure - just an honest discussion about what you're building and whether we're the right fit.