Synthesis Through Simulation: Enhancing Enterprise Data Generation Efficiency
The research introduces a new approach called **Synthesis Through Simulation** (STS), designed to improve how enterprises generate and evaluate data for AI applications, particularly in environments w...
Key Facts
- Enhance AI training efficiency by implementing Synthesis Through Simulation to generate diverse datasets.
- Reduce data access limitations by utilizing simulated environments for secure data generation processes.
- Boost operational flexibility by adopting STS, circumventing traditional database schema restrictions.
- Improve model accuracy by ensuring structurally valid data generation through controlled simulation practices.
- Integrate AI seamlessly into business operations by leveraging tool-calling agents enabled by STS.
Summary
Authors: Yipeng Li, Ashutosh Hathidara, Jane Lo, Harshavardhan Abichandani, Gunraj Singh, Atin Ghosh
Executive Summary
The research introduces a new approach called Synthesis Through Simulation (STS), designed to improve how enterprises generate and evaluate data for AI applications, particularly in environments where traditional methods face significant constraints. Companies often struggle with training AI models due to restrictions related to enterprise systems, data security, and database structures. These limitations hinder the effective use of tool-calling agents, which are crucial for integrating AI into business operations.
The STS framework overcomes these challenges by producing data through a simulated environment that adheres to operational policies. This simulation allows the generation of data without relying on specific database schemas, which often restrict access and flexibility. By executing operations in this controlled environment, STS ensures that the generated data is structurally valid, addressing a common problem in data synthesis where the authenticity and integrity of the data can be compromised.
A key component of STS is the Generalist Populator (GP), which operates without needing predefined database schemas. The GP is designed to enhance both the fidelity of the data generated and the scalability of the synthesis process. In tests across ten different simulated environments, the GP demonstrated an average marginal fidelity score of 0.88 and achieved 100% constraint satisfaction. This means that the data produced met all the necessary operational constraints consistently. In contrast, traditional statistical synthesizers struggled with seven of these environments due to their dependency on seed data, which is often unavailable in real-world scenarios. Similarly, agents that rely on specific schemas failed to complete 82% of tasks in a complex airline environment, highlighting the brittleness of schema-dependent approaches.
This research highlights the potential for STS to streamline data synthesis in enterprise settings, making it easier for companies to train AI models and integrate them into their operations. The ability to generate structurally valid data without the constraints of existing database schemas can lead to more robust AI applications. It may enable businesses to experiment with AI technologies more freely, ultimately accelerating the adoption of AI-driven solutions.
The research findings, including the STS framework and the ten simulated environments used for testing, have been made publicly available on GitHub. This open-source approach not only supports transparency but also encourages further innovation and development in the field of AI data synthesis.
Academic Abstract
Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained due to business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability, while procedure-based approaches yield the opposite weakness, typically lacking distributional fidelity without per-domain authoring. We introduce Synthesis Through Simulation (STS), a schema--free data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterprise environments. Because data is generated through the same environment that defines what is valid, STS guarantees structural validity by construction while decoupling validity enforcement from distribution modeling, allowing each to be addressed independently. The Generalist Populator (GP), STS's domain-agnostic agent, addresses the remaining challenges of distributional fidelity and synthesis scalability: GP achieves 0.88 average marginal fidelity and 100% constraint satisfaction across all ten environments without access to DB schemas, while statistical synthesizers are inapplicable to seven due to necessary seed data requirements, and schema-privileged agents fail 82% of trajectories on airline environment's tightly coupled workflows due to brittle task composition. We open-source the full framework, all ten environments, and generated datasets at https://github.com/SAP/synthesis-through-simulation.
Frequently Asked Questions
What business problems does this research solve?
This research addresses challenges related to generating and evaluating data for AI applications, particularly in situations where traditional methods are limited by enterprise system constraints, data security issues, and restrictive database structures.
Which industries benefit most from the Synthesis Through Simulation approach?
While the summary does not specify particular industries, any sector that relies on AI for data-driven decision-making and faces constraints in data generation and evaluation could potentially benefit from this approach.
What are the practical implementation considerations for the STS framework?
Implementing the STS framework may require consideration of the specific operational policies of the enterprise, as well as the setup of a simulated environment that mirrors real-world operations while allowing for flexible data generation.
What resources or expertise are needed to effectively utilize the Synthesis Through Simulation approach?
Organizations may need expertise in simulation technologies and a strong understanding of their operational policies, as well as resources to develop and maintain the simulated environments that generate the necessary data.
What are the competitive advantages of adopting the STS framework?
By enabling the generation of structurally valid data without relying on restrictive database schemas, the STS framework could provide enterprises with greater flexibility and agility in integrating AI into their operations, potentially leading to more informed decision-making and enhanced operational efficiency.