
How to Split a Multi-Document PDF Using JavaScript and Google Cloud Document AI
Introduction
In this tutorial, I will guide you through a process of splitting a PDF that contains multiple documents using JavaScript, Google Cloud’s Document AI, and the pdf-lib library. This feature is useful when you have a PDF with several documents, each identified by page numbers (e.g., “Page 1 of 2” for the first document, “Page 1 of 3” for the second document, etc.). Document AI will help extract page number data, and then we’ll split the PDF accordingly.
Step 1: Understanding the Problem
Consider a PDF with multiple documents, each identified by page numbers:
- The first document has 2 pages, labeled “Page 1 of 2”, “Page 2 of 2”.
The second document has 3 pages, labeled “Page 1 of 3”, “Page 2 of 3”, “Page 3 of 3”. We’ll use OCR (Optical Character Recognition) to extract these page numbers and split the PDF into separate files for each document.

Step 2: Setting Up Google Cloud Document AI
To OCR the page numbers, we will use Google Cloud Document AI’s Custom Extractor.
1. Create a Google Cloud Account if you don’t have one.
2. Set up Document AI by searching for it in the GCP Console

3. Create a Custom Processor by selecting the Custom Extractor model.

4. Select Custom extractor as our processor.

5. Upload Training Documents: Upload sample PDFs to train our processor .

6. Create Labels: Annotate the page numbers and total page count fields, creating two labels: page_no and page_total. For optimal accuracy, label at least 100 pages across 20 documents.


7. Train and Deploy the model.
Step 3: Extracting Page Numbers from the PDF Using Document AI
Once the processor is trained and deployed, you can extract labeled data like page numbers and total pages from the PDF. Here’s how we do it in JavaScript:
1const name = `projects/${projectId}/locations/${location}/processors/${processorId}`;2const buffer = await getTheArrayBufferFromPdfUrl(s3Url);3const encodedImage = Buffer.from(buffer).toString('base64');45const request = {6 name,7 rawDocument: {8 content: encodedImage,9 mimeType: 'application/pdf',10 },11};1213const [result] = await client.processDocument(request);14const { document } = result;15const { entities } = document;16const pages = formatData(entities);17const pagesToSplit = getPdfPagesToSplit(pages);
This function organizes the extracted data into a structured array containing each page’s number and total page count.
Step 4: Identifying Document Boundaries
We then determine the starting and ending pages for each document inside the PDF:
1getPdfPagesToSplit = (pages) => {2 const pdfPages = [];3 let count = 0;4 let skipCount = 0;56 for (const page of pages) {7 count++;8 if (skipCount) {9 skipCount--;10 continue;11 }1213 if (page.page_total == 1) {14 pdfPages.push({ number: +page.number + 1, start: count, end: count });15 } else if (page.page_total > 1) {16 skipCount = page.page_total - 1;17 pdfPages.push({ number: +page.number + 1, start: count, end: count + +page.page_total - 1 });18 }19 }2021 return pdfPages;22};2324 },25};2627const [result] = await client.processDocument(request);28const {document} = result;29const {entities} = document;30const pages = formatData(entities);31const pagesToSplit = getPdfPagesToSplit(pages);
Step 5: Splitting the PDF Using pdf-lib
Once we have the start and end pages, we can split the PDF using pdf-lib:
1extractPdfPage = async (arrayBuff, pageToSplit) => {2 const pdfSrcDoc = await PDFDocument.load(arrayBuff);3 const pdfNewDoc = await PDFDocument.create();4 const pages = await pdfNewDoc.copyPages(pdfSrcDoc, range(pageToSplit.start, pageToSplit.end));5 pages.forEach(page => pdfNewDoc.addPage(page));67 const newPdf = await pdfNewDoc.save();8 return newPdf;9};
Here, pdf-lib copies and saves the pages of each document as a new PDF.
Step 6: Upload or Download the Split PDFs
Now, we can take the split PDFs from SplittedPdfs and either upload them to a cloud service or download them to the user’s machine:
1const SplittedPdfs = [];2for (const pageToSplit of pagesToSplit) {3 const splittedPdf = await extractPdfPage(imageFile, pageToSplit);4 SplittedPdfs.push(splittedPdf);5}6// Now you can use SplittedPdfs as per your needs.
Conclusion
This tutorial demonstrates how to split a multi-document PDF using JavaScript, Document AI, and pdf-lib. We covered setting up Document AI, extracting page numbers, and splitting the PDF based on those page numbers. With these steps, you can easily implement this feature in your own applications.

Multi-Tenancy Patterns in DynamoDB: Silo, Pool, and Bridge Models
If you've already made the jump from a relational database to DynamoDB see our guide on moving relational data from SQL to DynamoDB...
Read More
We stopped leaving the IDE to design. Here’s our Cursor → Figma flow
Cursor drafts fast, catches gaps early, and still clips fields and breaks layouts. Here's the real pros-and-cons breakdown of our workflow...
Read More
AWS DevOps Agent: How AI is Automating On-Call Incident Response
If you've ever been on call during a production outage, you know how stressful it can be. Alerts start firing, dashboards light up, and suddenly you're jumping between monitoring tools...
Read More
Catch Missing Images Before Deploy: A Simple Pre-Build Script for Next.js
How Omax Tech added a lightweight image validation gate to Next.js 15 builds on Vercel...
Read More
Teach your LLM your design system: Storybook MCP + Amazon Bedrock + Strands
How to stop models inventing buttons and make them build UI from your real component catalog. Most "AI UI" demos look great until you paste the markup into a real product. The fix is not a smarter prompt...
Read More
Configure Self Hosted GitLab Repository Mirroring
Self hosted GitLab Repository Mirroring is a powerful feature that automatically synchronizes repositories between GitLab and external Git providers...
Read More
Event Sourcing: A Foundation Guide
Event Sourcing is an architectural pattern where every state change is recorded as an immutable event rather than updating a database row in place...
Read More
AWS Security Best Practices Every Business Should Follow
As more organizations migrate their applications and critical workloads to AWS, securing cloud environments has become a business priority rather than just an IT responsibility...
Read More
AWS Migration Checklist: A Practical Roadmap for Modern Businesses
Migrating businesses to AWS offers many benefits, including cost optimization, improved security, and greater scalability. However, a successful migration requires careful planning and execution. Otherwise, organizations may experience...
Read MoreReady to Work With Us?
Most engagements start with a 20-minute conversation. No pitch, no pressure - just an honest discussion about what you're building and whether we're the right fit.