GroupDocs.Parser at a glance

Document Parser SDK for performing high‑accuracy document parsing in Node.js applications

Illustration parser

Extract data from documents

GroupDocs.Parser for Node.js via Java API enables you to retrieve text, metadata, and images from a wide range of file formats such as Office documents, emails, attachments, and archives. This powerful tool helps you efficiently access and process valuable information contained within these files for various applications like data analysis, search engine indexing, or content management systems.

Parse documents

Extract various elements such as hyperlinks, tables, QR codes, barcodes and data from PDF forms. Also parse any desired information from documents using custom templates.

Customizing results

Node.js API enables you to retrieve data in various formats such as raw, structured, HTML, or Markdown. Additionally, the API offers a search functionality for locating specific words or phrases within the text of documents.

Platform Independence

GroupDocs.Parser for Node.js via Java supports the following operating systems, frameworks and package managers

Windows
macOS
Linux
NPM
Amazon
Docker
Azure
VS Code
IntelliJ

Supported file formats

GroupDocs.Parser for Node.js via Java supports operations with the following file formats.

Microsoft Office formats

  • Word: DOCX, DOC, DOCM, DOT, DOTX, DOTM, RTF
  • Excel: XLSX, XLS, XLSM, XLSB, XLTM, XLT, XLTM, XLTX, XLAM, SXC, SpreadsheetML
  • PowerPoint: PPT, PPTX, PPS, PPSX, PPSM, POT, POTM, POTX, PPTM

Images & Other Formats

  • Portable: PDF
  • Images: JPG, BMP, PNG, TIFF, GIF
  • Other office formats: ODT, OTT, OTS, ODS, ODP, OTP, ODG

Other formats

  • Web: HTML, MHTML
  • Archives: ZIP, TAR, 7Z
  • e-Books: CHM, EPUB, FB2, MOBI

GroupDocs.Parser for Node.js via Java features

Extract data from PDFs, Office documents, images and other formats swiftly and accurately with our Node.js Document Parser SDK

Feature icon

Extract text

Extract textual information from various file formats such as office documents, PDF files and images for easy readability and analysis.

Feature icon

Extract images

Retrieve visual content from diverse sources like office documents, PDF files for convenient access and use.

Feature icon

Scan QR Codes

Detect and decode QR codes present within office documents, PDF files, or visual content for efficient information retrieval.

Feature icon

Extract data from email attachments and archives

Gather valuable information from email messages, file attachments, and compressed data sources for effective analysis and utilization.

Feature icon

Extract tables

Identify and extract tabular data from PDF documents for organized analysis and use.

Feature icon

Extract hyperlinks

Locate and extract hyperlinks and email addresses within office documents or PDF files for efficient access.

Feature icon

Parse PDF Forms

PDF Forms are digital documents featuring fillable fields for user interaction, allowing them to input information electronically. Node.js API can be utilized to extract data from these forms for efficient processing.

Feature icon

Parse data by templates

Create custom templates and utilize them with Node.js API to parse specific information from PDF files, simplifying data extraction processes.

Feature icon

Search a text in documents

Quickly locate specific words or patterns within documents.

Code samples

Beyond basic text extraction, here are the most common use cases for quick text, image and metadata extraction.

Search Text in a Document

This example shows how to search for a word in a PDF document and print the page and position of every match.

Search Text in a Document in JavaScript

const groupdocs = require('@groupdocs/groupdocs.parser');

// Load the document
const parser = new groupdocs.Parser("sample.pdf");

// Search for a word page by page
const options = new groupdocs.SearchOptions(false, false, false, true);
const results = parser.search("Lorem", options);

// Print the page index, position and text of every match
const it = results.iterator();
while (it.hasNext()) {
    const result = it.next();
    console.log(`Page ${result.getPageIndex()}, position ${result.getPosition()}: ${result.getText()}`);
}

parser.close();

Extract Images from a Document

This example shows how to extract images from a PDF document and save them as PNG files.

Extract Images from a Document in JavaScript

const groupdocs = require('@groupdocs/groupdocs.parser');

// Load the document
const parser = new groupdocs.Parser("images.pdf");

// Extract images from the document
const images = parser.getImages();

// Save the images as PNG files
const options = new groupdocs.ImageOptions(groupdocs.ImageFormat.Png);
let index = 1;
const it = images.iterator();
while (it.hasNext()) {
    it.next().save(`image_${index++}.png`, options);
}

parser.close();

Extract Metadata from a Document

This example shows how to extract metadata from a PDF document and print it.

Extract Metadata from a Document in JavaScript

const groupdocs = require('@groupdocs/groupdocs.parser');

// Load the document
const parser = new groupdocs.Parser("sample.pdf");

// Extract metadata from the document
const metadata = parser.getMetadata();

// Print the metadata
const it = metadata.iterator();
while (it.hasNext()) {
    const item = it.next();
    console.log(`${item.getName()}: ${item.getValue()}`);
}

parser.close();

Ready to get started?

Download GroupDocs.Parser for free or get a trial license for full access!

Useful resources

Explore documentation, code samples, and community support to enhance your experience.

Temporary license tips

1
Sign up with your work email.
Free mail services are not allowed.
2
Use Get a temporary license button on the second step.
✕
 English