Skip to main content

Extract data from PDF and text documents

Learn how to extract structured data from PDF and Google Docs files and prepare the results for merging into one table.

Written by Jonatan Gomes

Sheetgo can extract data from the documents in a folder, including PDF, Google Docs, Word, and plain text files. The automation uses AI to turn the documents into structured rows in a single Google Sheets table.

Create a document extraction automation

  1. Open your workflow and select Add to workflow > Automation.

  2. Choose Folder as the source.

  3. In the Folder field, click Select folder and choose the folder that contains your documents. After selection, use Change folder to choose a different folder.

  4. Under IMPORT SETTINGS, open Content type and select Documents. This removes the tab picker and displays IMPORT OPTIONS.

The default Settings option is Import data from every file in the folder (newest to oldest).

Choose which documents to import

Filter folder content is optional. Leave it off and Sheetgo imports every file in the folder. Turn it on when you want only some of the files, then set Match to all or any and build your conditions.

The criteria and the value control depend on the field you choose:

Field

Criteria

Value

File type

is, is not

Choose PDF, Document, or Image from the list

File name

contains, does not contain, starts with, is

Text

Date modified

modified within, modified before, modified after, between

Pick a date

File content

is not empty, is not near-empty, is empty, is near-empty

No value. The input is disabled.

Select Add condition for another condition. Select Add group to create nested logic.

Under IMPORT OPTIONS, Only import each file once (skip already-processed files) is on by default in Documents mode. Keep it on when you want later runs to process new files without importing processed files again. If a rerun finds no new files, it adds nothing.

The source step also has these options under FOLDER OPTIONS in the advanced settings:

  • Include files in subfolders

Tell the AI what to extract

The second step is Process with AI, marked BETA. AI processing is required for documents. This step extracts data from each document into structured rows in your spreadsheet.

Enter the fields and row structure you need in the INSTRUCTIONS box. You can also start with one of the EXTRACTION PRESETS:

  • Invoices

  • Receipts

  • Contracts

  • Other documents

For example, the Invoices preset fills the instructions with:

Extract each invoice as a row with invoice number, date, supplier, total and tax.

This instruction creates one row per invoice with the columns invoice number, date, supplier, total, and tax.

To create one row per line item instead, tell the AI to extract every line item as one row. A document with three line items then produces three rows. Fields that do not appear in a document remain empty; the AI does not invent values.

Use exact column names

Name every column explicitly, and tell the AI to keep the names and order unchanged. For example:

Extract every line item as one row. Use exactly these column headers, in this order: invoice number, date, supplier, item, quantity, unit price, line total. Do not add, rename, or reorder columns.

This matters because vague instructions can produce different headings from one document to the next, such as Invoice #, Invoice Number, and InvoiceNumber. Naming the columns keeps every document writing into the same columns of the table.

Set how processing failures are handled

Open the AI step's advanced settings. Under EXTRACTION OPTIONS, set When a document fails to process to Continue with the other files if you do not want one failed document to stop the others. Sheetgo will notify you about files it could not process.

The AI step also offers Connect your OpenAI account for unlimited AI processing. This option uses your key to remove token and data size limits.

Choose the destination

Choose Google Sheets as the destination, then select New file or Existing file. For a new file, complete the File name field. The available FILE TYPE choices are Google Sheets, Excel, and CSV.

Understand the extraction output

The automation writes all of the extracted data into a single tab in the destination file. Each document becomes one row, or several rows if your instructions ask for one row per line item.

The columns come from your instructions or from the preset you chose. Every document writes into the same columns, so three invoices processed with the Invoices preset produce a table like this:

invoice number

date

supplier

total

tax

INV-2026-0143

August 4, 2026

Vantage Office Supplies

£626.10

£104.35

INV-4471

August 11, 2026

Riverstone Logistics

£958.08

£159.68

2026-0087

August 19, 2026

Corvus Print Studio

£1,822.80

£303.80

Extracted dates arrive as real dates, so you can sort and filter the column without converting it first.

Keep reading: to process each new file within minutes of it arriving in the folder, see Run a workflow when a file or folder changes.

Did this answer your question?