How to Automate Bill of Lading Data Extraction (No Custom Parser Needed)

The BOL bottleneck

Most bills of lading you handle today still arrive as a scan, a PDF, or a fax, not as structured data. According to the Digital Container Shipping Association, electronic bills of lading made up just 11% of the roughly 45 million bills issued industry-wide by mid-2025, up from about 1% in 2021. The largest ocean carriers, Maersk, MSC, CMA CGM, and Hapag-Lloyd among them, have pledged full electronic adoption, but not until 2030.

McKinsey estimates that universal eBL adoption could save the industry $6.5 billion a year in direct documentation costs, on top of $30 to $40 billion in broader trade growth from removing that friction. That's the size of the problem sitting inside one trade document, and it isn't closing for at least another five years.

Which means the bill of lading landing in your inbox this week showed up however that carrier's system happened to produce it. Ocean, domestic straight, and multimodal are three different document families before you even account for variation within each carrier. Shipper and consignee fields swap position between forms. Container and commodity data show up as a clean table on one BOL and a run-on paragraph on the next.

Waiting for an industry-wide standard to solve this isn't a plan your operations team can run on. Handling the document you actually received, however it's formatted, is.

What actually needs to come out of a BOL

A bill of lading carries more than a shipment description. Legally, it's doing three jobs at once: it's a receipt for the goods, evidence of the contract of carriage, and in a negotiable form, a document of title. Whoever holds the original can claim the cargo. That's part of why the format resists standardization, carriers built their BOLs around slightly different legal and operational priorities, but the core data your systems need out of it stays consistent.

Field group What's in it
Parties Shipper, consignee, notify party
Ports and routing Port of loading, port of discharge, vessel and voyage number
Container and commodity Container numbers, seal numbers, commodity description, HS codes
Weight and packaging Gross weight, volume, package count and type
Commercial terms Freight charges, payment terms (prepaid or collect), special instructions (hazmat, temperature control)

Every one of these fields shows up on nearly every BOL you'll process. What changes is where it sits on the page, what it's labeled, and whether it appears as a tagged field, a table row, or a sentence buried in a shipping instructions block.

That's the part a fixed template can't absorb. Which raises the real question: whether "parsing" a bill of lading is even the right way to think about this problem in the first place.

Why "parser" is the wrong mental model

A parser, in the traditional sense, is built around a fixed layout. You tell it where the shipper's name sits on the page, or write a rule that looks for a specific label, and it works exactly as long as every document that comes through matches that layout. That's how most legacy freight-document tools were built, and it's why so many freight forwarders still keep a folder of carrier-specific templates someone has to update every time a carrier tweaks their form.

Bills of lading break that model constantly. An ocean BOL from one carrier and a domestic straight bill from a regional trucking company don't share a layout, a field order, or sometimes even field names. Even the same carrier's BOL shifts between their standard form and a hazmat variant. A parser tuned to last month's format fails silently on this month's revision, and nobody notices until a shipment gets held up because the consignee field came through blank.

Foxello's approach starts from a different question: not "where does this field sit," but "what does this field mean." You describe what you're looking for, shipper name, container number, commodity description, once, in plain language, when you set up the workflow. The extraction model reads each document for that meaning wherever it appears, rather than matching it to a fixed position. A container number is still a container number whether it's in a labeled field, a table cell, or a line of shipping instructions.

That shift matters most once you get to the messiest part of a BOL: the commodity and line-item table itself.

Commodity and line-item data

The commodity section of a bill of lading is where automation earns its keep, and where it's hardest to get right. A single BOL might list one container with one commodity, or a dozen line items across multiple containers, each with its own weight, package count, and commodity description. Some carriers lay this out as a clean table. Others write it as a paragraph of shipping instructions with the same information buried in prose.

Foxello handles this the same way it handles invoice line items or claims-history rows: as a table-type field. You define the columns once, container number, commodity description, weight, package count, and every row across every BOL variant gets pulled into that same structure, whether the source document presented it as a table or not. Rows get added or removed automatically to match what's actually on the page, not what a template expects.

Getting this right matters past extraction. Inbound Logistics' analysis of LTL billing disputes ties a meaningful share of freight invoice errors directly to commodity description and classification mistakes introduced during manual transcription, the kind that trigger a reweigh, a reclassification, or a dispute with the carrier over the final bill. Structured, accurate commodity data at the extraction stage is what keeps that downstream invoice matching clean.

That table-level accuracy is also the foundation for the question customers ask most: whether they can define their own fields for a BOL variant nobody else uses.

Custom schema definitions, answered directly

If you're searching for a platform that lets you define your own schema for bill of lading processing, here's what that actually needs to do: let you list the fields you want, in your own words, without touching a template, a regex, or a developer.

In Foxello, that's the Instinct or Mastery workflow setup. You describe each field once, shipper name, container number, commodity description, whatever your operation tracks, in the workflow editor. That's a one-time setup per workflow, not something you repeat per file or rebuild every time a carrier updates their form. Instinct suits a new carrier or a shipment type you haven't standardized around yet; Mastery earns its keep once you're running high volume against a format that stays reasonably consistent, like a single carrier or lane you handle daily.

This also answers the harder version of the question: what happens when your BOLs aren't all the same. If two carriers structure their documents differently enough that one workflow can't serve both well, you run two workflows, each with its own field list, rather than forcing one schema to fit documents that were never built to match. There's no schema file to hand off to engineering and no vendor ticket to open when a carrier changes their layout. You edit the field list yourself, in the same editor where you set it up.

That's the actual differentiator worth checking when you're evaluating platforms: not whether a vendor supports bills of lading, most say they do, but whether defining what you need is something your own team can do in minutes, or something that goes into someone else's backlog.

What good BOL automation looks like operationally

Once the fields and schema are defined, the workflow around them is what actually determines whether a team stays fast at volume. This is where import channel choice matters as much as extraction accuracy.

BOLs arrive from wherever your carriers and forwarders send them, and email is the most common intake point in logistics, since brokers and carrier reps still send scanned or PDF bills as attachments. Foxello's email import matches each incoming message to its workflow using a generated import address, filters by sender and attachment type if you've configured either, and creates a file automatically without anyone forwarding or uploading it by hand. If your TMS or forwarding platform can push files directly, API import skips the inbox step entirely and lets a shipment system post BOLs the moment they're received.

From there, processing runs without anyone touching it: OCR, extraction against the fields you defined, and a confidence score on every value. That's where review setup earns its keep. Confidence-based review holds only the fields that come back below the confidence threshold for a person to check, a container number that scanned ambiguously, a commodity description split across a page break, while everything else auto-confirms and moves on. High-volume operations don't get faster by reviewing every field on every BOL. They get faster by only reviewing the field that actually needed a second look.

Export closes the loop. JSON or CSV output maps cleanly into most TMS and ERP systems for reconciliation against the shipment record, and webhook or Zapier/n8n export means the structured data lands where your team already works instead of sitting in a downloads folder waiting for someone to move it.

That's the operational shape of it: documents arrive without manual upload, extraction happens without a template, review happens by exception, and the data lands in the system that needs it, not back in someone's inbox.

Getting started checklist

None of this requires new infrastructure. It requires one workflow, set up once, per BOL type you handle regularly.

To set it up:

  • Create a workflow and choose Instinct if you're handling a new carrier or shipment type, Mastery if you're running high volume against a format that stays consistent.
  • Describe the fields you need in plain language: shipper, consignee, ports, container and commodity data, weight, freight terms. One-time setup, not a per-carrier template.
  • Define the commodity table columns you need once, container number, description, weight, package count, so every BOL's line items land in the same structure regardless of how the source document laid them out.
  • Set the import channel to match how BOLs actually arrive: email for broker and carrier-sent attachments, API if your TMS can push files directly.
  • Choose confidence-based review so only uncertain fields, an ambiguous scan, a split commodity description, land in front of a person.

Either way:

  • Start with your highest-volume carrier or lane first. That's where the current manual process is costing the most hours, and where the gain shows up fastest.
  • If a second carrier's BOL doesn't fit the same field list well, run it as its own workflow rather than stretching one schema to cover documents that were never built to match.

Electronic bills of lading are still years from being the default. Until they are, the practical fix is handling the paper and PDF BOLs you actually receive, however they're formatted, without waiting on the industry to standardize first. Start a workflow, describe your fields, and run your next BOL through it before deciding anything further.

Ready to automate your documents?

Start processing your first documents in minutes. No setup required.

Start free - no card