For many enterprises, the promise of artificial intelligence has shifted from a futuristic vision to a pressing operational reality. However, the transition from a successful pilot project to a scalable production environment often reveals a “messy truth”: the infrastructure required to support AI is frequently fragmented, inefficient, and difficult to maintain. This gap between strategic ambition and technical execution is where many companies find themselves struggling with what experts call pipeline sprawl and the emergence of “shadow AI.”
The challenge lies in the nature of enterprise data. While organizations possess vast amounts of relational data—tracking customers, transactions, and behaviors—turning that information into accurate, actionable predictions is rarely a seamless process. Traditionally, this has required months of manual effort by specialized data science teams to build and maintain complex feature engineering pipelines, creating a bottleneck that slows down the time-to-value for AI initiatives.
Addressing these inefficiencies requires a shift in how companies approach the “plumbing” of AI. By focusing on reducing complexity and automating the way data is processed, enterprises can move away from fragile, manual workflows toward more robust systems. Dr. Hema Raghavan, co-founder of Kumo.AI, has spent her career tackling these specific hurdles, moving from leading machine learning teams at LinkedIn to building tools designed to democratize sophisticated AI techniques for the broader enterprise market.
The goal for modern AI strategies is no longer just about the power of the model, but about the accessibility of the system. When the interface is too complex, the technology remains siloed within a small group of experts. When the infrastructure is too costly, the ROI becomes impossible to justify. Solving the “messy” side of AI implementation means building simple interfaces for complex technology and optimizing infrastructure for real-world economics.
Overcoming the Complexity of Enterprise Data Pipelines
A primary driver of “pipeline sprawl” is the manual labor involved in feature engineering. In a typical enterprise setting, data is stored in relational databases like Snowflake or Databricks. To make this data useful for a machine learning model, data scientists must manually extract, transform, and load (ETL) the data, often creating a series of fragile pipelines that break whenever the underlying data schema changes.
This manual process creates a significant barrier to entry. According to Kumo.AI, turning relational data into accurate predictions often requires months of work by specialized teams. This delay not only hinders innovation but also increases the likelihood of “shadow AI,” where business units deploy unauthorized, fragmented AI tools to bypass the slow official pipelines of the IT department.
To combat this, the industry is seeing a move toward “foundation models” for structured data. Kumo developed KumoRFM, a foundation model specifically designed for structured enterprise data. This approach allows organizations to generate predictions directly from their data warehouse without the need for task-specific model training, effectively removing the manual feature engineering bottleneck and reducing the sprawl of disparate data pipelines.
The Role of Graph Neural Networks (GNNs)
One of the most effective ways to handle the complexity of relational data is through Graph Neural Networks (GNNs). Unlike traditional models that treat data points as isolated entries, GNNs can capture the relationships and connections between data points—such as how a customer interacts with multiple products over time or how a transaction relates to a network of other users.
Historically, GNNs were considered too complex for most enterprises to implement since they required specialized expertise and significant computational resources. However, by creating SQL-like query languages, companies can now make GNNs accessible to non-technical users while still providing full control to data scientists. This democratization allows a broader range of employees to derive insights from their data without needing a PhD in machine learning.
Balancing Performance and Economic Reality
The “messy truth” of AI strategies also extends to the cost of compute. There is a common tendency in the AI industry to chase maximum performance at any cost, often relying exclusively on high-end GPUs. While GPUs provide breakthroughs in performance, they can be prohibitively expensive for some enterprise-scale applications, leading to a gap between a model’s theoretical capability and its economic viability.

A more sustainable approach involves optimizing infrastructure for real-world economics. For example, implementing a hybrid CPU/GPU approach can dramatically improve cost efficiency without sacrificing the capabilities of the AI. This thoughtful design ensures that AI is not just a research experiment but a scalable business tool that provides a clear return on investment (ROI).
The focus must shift toward “time-to-value.” The faster a customer can see a result from their data, the more likely they are to expand their usage of the technology. By removing the requirement for customers to manually convert their data into graph formats or build complex pipelines, companies can focus on business outcomes rather than technical maintenance.
From LinkedIn Scale to Enterprise Innovation
The transition from managing AI at a massive scale to building AI tools for others requires a deep understanding of how systems fail as they grow. Dr. Hema Raghavan’s experience leading machine learning teams at LinkedIn—during a period where the platform grew from 400 million to 700 million users—provided critical insights into the necessity of reducing complexity in AI systems.

At that scale, systems like “People You May Know” require immense stability and efficiency. Applying those lessons to the enterprise market, Raghavan recognized that the same principles of scalability and simplicity are needed for businesses trying to operationalize AI. This led to the founding of Kumo in 2021, with a mission to make the power of graph neural networks accessible to any company with a data warehouse.
This commitment to innovation has been recognized externally. Dr. Raghavan was named to Inc.’s 2026 Female Founders 500 list, which recognizes impactful women entrepreneurs shaping the future of their industries. This recognition underscores the growing importance of leadership that prioritizes both technical excellence and business pragmatism in the AI era.
Key Takeaways for AI Strategy
- Reduce Pipeline Sprawl: Move away from manual, task-specific feature engineering toward automated solutions and foundation models for structured data.
- Democratize Access: Implement simple, intuitive interfaces (such as SQL-like languages) that allow non-technical users to leverage complex AI capabilities.
- Optimize for ROI: Avoid “performance at any cost.” Utilize hybrid infrastructure (CPU/GPU) to balance computational power with cost efficiency.
- Prioritize Time-to-Value: Focus on tools that allow users to generate predictions directly from existing data warehouses rather than requiring extensive data reformatting.
As enterprises continue to refine their AI strategies, the focus will likely shift further toward operationalization—moving beyond the “wow” factor of generative AI to the “work” of predictive AI that drives actual revenue and efficiency. The next critical checkpoint for many organizations will be the transition to fully automated data pipelines that can adapt to changing business needs in real-time.
Do you struggle with “pipeline sprawl” or the challenges of operationalizing AI in your organization? Share your experiences in the comments below.
Worth a look