Processing Unstructured Documents - Semantic
Processing Unstructured Documents

Processing Unstructured Documents

2025.07.23

In our view, there are three different approaches worth considering when looking for a solution to process unstructured documents:

- Cloud-based document processing services
- Proprietary large language models (LLMs)
- Open-source large language models (LLMs)

Cloud-based document processing services

Cloud providers have long offered specialized services for extracting data from documents:

  • AWS Textract
  • Azure AI Document Intelligence
  • Google Document AI

These solutions offer similar capabilities, including OCR, table extraction, filters, and many other features. They also support a wide range of document formats out of the box. They can be integrated into existing systems through APIs, which means that the source documents must be sent to the cloud for processing.

Advantages:

  • Quick and easy to get started
  • High accuracy
  • Scalability
  • No need for local server infrastructure

Disadvantages:

  • Vendor lock-in
  • Limited flexibility
  • Data privacy concerns

Proprietary large language models (LLMs)

The best-known AI providers make their large language models available through APIs. With this approach, the document must first be prepared and then sent through an API for processing within the provider’s infrastructure.

Proprietary LLMs include:

  • OpenAI GPT
  • Anthropic Claude
  • Google Gemini
  • Mistral AI

We recommend OpenAI, as we currently consider it the most advanced solution on the market. OpenAI uses Microsoft Azure data centers, offers a zero-data-retention option, and can also provide processing within EU data centers upon request.

Advantages:

  • Flexibility
  • Better contextual understanding
  • Rapid model development
  • Multimodal capabilities, supporting both text and image inputs
  • Additional features can be implemented, such as customer service chatbots
  • No need for local server infrastructure

Disadvantages:

  • Data privacy concerns
  • Vendor lock-in, although many other LLMs support OpenAI-compatible APIs
  • Operates as a “black box,” with no transparency into its internal processes

Open-source large language models (LLMs)

Open-source large language models are generative AI solutions that can be deployed and trained locally. Their base models undergo extensive pre-training, enabling them to understand context. They can also be further trained for specific domains and use cases.

Open-source and open-weight LLMs include:

  • Meta Llama
  • Google Gemma

We recommend Meta’s Llama model, which is developing rapidly and also offers multimodal capabilities. For example, Llama 3.2 can process images as well as text.

Advantages:

  • No data privacy concerns, as documents remain within your own infrastructure
  • No vendor lock-in
  • Can be developed into a highly customized solution through long-term model training
  • Greater transparency
  • Complete control

Disadvantages:

  • Higher initial costs due to the need for local server infrastructure
  • Accuracy depends on model size, with larger models requiring substantial computing capacity
  • Scalability challenges
  • Server capacity may not be used efficiently

Summary

It is difficult to predict in advance how well an AI model will perform on a particular task. Models are extremely complex and non-deterministic, while the input PDF documents themselves are often highly varied: they may be scanned and use very different layouts and formatting.

It is also important to remember that this field is evolving rapidly. For this reason, it is best to avoid becoming dependent on a single provider or solution.

The system should be designed so that the language model can be replaced easily. This makes it simpler to take advantage of new model versions and more advanced solutions while reducing vendor lock-in.

Author

András Kenéz
Head of Development, Software Architect
December 5, 2024