spot_img
HomeResearch & DevelopmentFlexDoc: A New Approach to Generating Diverse Synthetic Documents...

FlexDoc: A New Approach to Generating Diverse Synthetic Documents for AI Model Training

TLDR: FlexDoc is a novel framework that uses stochastic schemas and parameterized sampling to create diverse, multilingual, and richly annotated synthetic documents. This approach significantly reduces the cost and effort of data collection for training document understanding models, improving their performance on tasks like Key Information Extraction by augmenting real datasets, and offering a scalable alternative to traditional hard-template methods.

Developing advanced AI models for understanding documents is a crucial task for many businesses, from processing invoices to verifying identities. However, a major hurdle in this area is the need for vast amounts of diverse, well-annotated data. Collecting such data is not only incredibly expensive, potentially costing millions of dollars, but also faces significant challenges related to privacy and legal restrictions.

A new research paper, titled “FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models,” introduces an innovative solution to this problem. Authored by Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal, Ranjeet Gupta, Amit Agarwal, Praneet Pabolu, Srikant Panda, Hansa Meghwani, Graham Horwood, and Fahad Shah from Oracle AI, the paper details FlexDoc, a scalable framework designed to generate realistic, multilingual, and semi-structured synthetic documents with rich annotations.

FlexDoc tackles the limitations of traditional data collection by probabilistically modeling various aspects of documents, including layout patterns, visual structure, and content variability. This allows for the controlled generation of a wide array of document variants at scale, effectively bypassing the high costs and privacy concerns associated with real-world data.

How FlexDoc Works

The core of FlexDoc lies in its three main components: Stochastic Schemas, a Parameterized Sampling algorithm, and a Document Rendering algorithm.

Stochastic Schemas: Instead of rigid templates, FlexDoc uses “stochastic schemas” where document elements and their properties are defined as random variables. This means that attributes like the number of rows in a table or the presence of a specific entity can be sampled from defined distributions or probability ranges. This flexibility allows a single schema to generate hundreds of thousands of unique document permutations.

Parameterized Sampling: This algorithm takes the stochastic definitions from the schema and “freezes” them for each document instance. It generates fake values for entities (like names, addresses, or dates) using a type-specific generator, ensuring privacy and diversity. It also samples layout and style attributes, creating a unique “document permutation” for each generated document.

Dynamic Virtual Grid Algorithm: To ensure visual diversity and prevent overlaps, FlexDoc employs a clever “Dynamic Virtual Grid Algorithm.” This system treats each document section as a virtual grid, dynamically adjusting cell sizes to place entity groups (like ‘Merchant Details’ or ‘Invoice Details’) without overlap, while maintaining visual structure.

Document Rendering: Finally, the Document Rendering algorithm draws these sampled entity groups onto a blank canvas, recording precise bounding boxes and class labels for annotation. This output includes both the rendered image and the detailed annotations needed for training AI models.

Multilingual Capabilities

One of FlexDoc’s standout features is its multilingual support. A single stochastic schema, initially defined in English, can be used to generate documents in different languages. By simply switching a configuration for the fake value generator and enabling machine translation for headers, FlexDoc can produce documents with locale-specific values and translated text, making it highly adaptable for global enterprise applications.

Impressive Results

Experiments on Key Information Extraction (KIE) tasks, a complex aspect of document understanding, demonstrated FlexDoc’s effectiveness. When FlexDoc-generated data was used to augment real datasets, it improved the absolute F1 Score by up to 11% for models like LayoutLM and Phi-4-Multimodal. This performance boost was observed in both English (DocILE dataset) and Spanish (IDSEM dataset) invoices, highlighting its potential for multilingual applications.

Furthermore, FlexDoc significantly reduces annotation effort. Compared to traditional hard-template methods, it achieved comparable model performance while cutting annotation time by over 90%. This means that the cost of generating 3,000 samples with FlexDoc remains constant at about 150 minutes, whereas hard-template methods require a linear increase in time as the dataset grows.

The framework also generates diverse data, with similarity levels comparable to real-world benchmarks, ensuring that models trained on synthetic data are robust and generalize well.

Also Read:

Future Directions and Limitations

While FlexDoc is highly effective for semi-structured documents like invoices and receipts, the authors acknowledge its limitations for fully structured, visually rich documents such as driving licenses or passports. These documents typically have limited layout variations, and hard-template approaches might still be more suitable for them, especially when visual backgrounds are critical.

Future work aims to extend FlexDoc to address these limitations, potentially by integrating semantic value generators to ensure logical consistency (e.g., invoice totals adding up correctly) and by accounting for cultural variations in document generation across different languages. You can read the full research paper here: FlexDoc Research Paper.

FlexDoc is already in active deployment, accelerating the development of enterprise-grade document understanding models and significantly reducing data acquisition and annotation costs, proving to be a robust and extensible framework for the future of AI in document processing.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -