HyT Capital Portfolio | KeenData: Building an Integrated Data & AI Architecture, Constructing an AI Data Foundation Through "Model-Data Resonance"
The following article is sourced from "DataYuan" by Yue Man Xi Lou (月满西楼).
China Doesn't Need the 101st Open-Source LLM — It Needs a Stronger AI Data Infrastructure
June 2026, Moscone Center, San Francisco. The Databricks Data+AI Summit drew 30,000 attendees and featured over 18 major product announcements. CEO Ali Ghodsi said something on stage that was widely quoted: "This is going to get extremely expensive — and we're just getting started."
July, Beijing National Convention Center. At the Global Digital Economy Conference, one Chinese company's booth was surrounded by a crowd.
Between these two events lies the Pacific Ocean — and an assumption that is being validated in real time: the second half of AI is shifting direction.
Over the past two years, large language models have gone from hundreds of billions of parameters to trillions. Open-source and closed-source models have taken turns topping the leaderboards. Companies have been swept up in the frenzy — buying compute, building teams, deploying models — convinced they had secured a ticket to the future.
The result?
Models are "omnipotent" in general Q&A, but as soon as they enter factories, hospitals, or energy sites, they fail to adapt. Data is scattered across dozens of silos — inconsistent formats, mismatched standards, uneven quality. The data simply cannot be "fed" into the models.
At the 2026 Global Digital Economy Conference, the focus of discussion began to shift. Beyond asking "which model has the highest benchmark score," people started asking a more practical question: how can AI truly be deployed on the ground?
A consensus is forming: the second half of AI will be decided not just by model benchmarks, but by AI data infrastructure.
The AI Industry Is Switching Tracks
From "Competing on Models" to "Competing on Data Foundations"
Let's start with a typical scenario.
The CIO of a large state-owned enterprise did the math: the group has 17 business systems, 9 data warehouses, and 3 cloud platforms, with data formats ranging from Oracle to MySQL to Hadoop to unstructured documents — a chaotic mix. The group mandated a "full embrace of AI," so they purchased large model platforms, built computing clusters, and assembled an algorithm team.
Six months later, the model was deployed — and it couldn't even run.
"We wanted to do something very simple — using a large model to assist with equipment fault diagnosis," the CIO said. "But the model couldn't even understand the equipment's historical maintenance records, because the records were scattered across three systems, with different formats, different field names, and a large number of paper report scans that hadn't been processed."
This is not an isolated case.
The real pain point of enterprise AI deployment: money spent, but the model won't run.
Data scattered across dozens of silos — inconsistent formats, mixed standards, poor quality — is the reality for the vast majority of enterprises. Traditional sectors like energy, manufacturing, and healthcare have accumulated petabytes of data, but most of it lies dormant in isolated systems, difficult to be directly utilized by algorithms.
Large models are powerful, but if they are given "unwashed vegetables and an unorganized warehouse," they cannot produce a good meal.
To some extent, the bottleneck for AI at scale has shifted from "whether we can compute" to "whether the data is usable."
It's not that the algorithms aren't good enough, or the compute isn't powerful enough — it's that the data system, governance across the entire data landscape, and the full-chain engineering closed-loop capability can't keep up. An enterprise that cannot continuously supply high-quality data is like a kitchen without ingredients — no matter how advanced the equipment, it cannot produce a good dish.
Yu Yang, Chairman of KeenData, put it bluntly: "In the first half of AI, the industry focused more on building general-purpose capabilities. In the second half, when AI needs to deliver greater impact for large enterprises, industries, and government entities, we must solve the last-mile problem: data."
Databricks and Snowflake Both Pivot, Aligning in the Same Direction
If the confusion of Chinese enterprises were still just an isolated phenomenon, then the collective moves of the global AI industry chain signal the start of a new industry cycle.
Let's look at the model side first.
Anthropic conducted a controlled experiment: the same Claude model, when run naked on enterprise data analysis tasks, achieved only 21% accuracy. After integrating a complete data engineering system, accuracy jumped to over 95%. The gap was not in the model itself, but in the data infrastructure behind it.
Anthropic subsequently built a four-layer architecture: "data infrastructure layer → source of truth layer → skill layer → validation closed-loop layer." This architecture has now become a common reference paradigm for both Databricks and Snowflake.
Now look at the data platform side.
Since 2025, Anthropic has established deep partnerships with Databricks and Snowflake. Claude is natively integrated into the Databricks Data Intelligence Platform and interconnected with Unity Catalog, enabling unified permissions, auditing, and data staying within the boundary. At the same time, it is deeply embedded into Snowflake Cortex AI, completing natural language queries, code generation, and AI agent development within the data boundary.
Model vendors and data platform vendors are moving from "going it alone" to "deep integration."
At the Databricks Data+AI Summit in June 2026, the results of this convergence were on full display. Databricks redrew its entire architecture diagram: Lakehouse moved down to the foundation layer, Agent Runtime rose to the top, and governance was repositioned as the first checkpoint before AI deployment.
The signal was clear: models and data infrastructure are forming a mutually reinforcing relationship. Models depend on the data foundation to supply "understandable and trustworthy" data; the data foundation, in turn, leverages the power of models to activate dormant data assets. They are no longer in a linear upstream-downstream relationship, but a closed loop of mutual reinforcement.
Snowflake is doing the same. The two long-time competitors in the cloud data platform space have reached a rare consensus on this point: the next round is about "whether the data is ready."
How Should China's AI Data Infrastructure Be Built?
From San Francisco back to Beijing — shifting the perspective from global to local — the question is essentially the same: as AI and Agents move from technology competition to industrial deployment, is our data infrastructure ready?
China's industrial scenarios are more complex. Manufacturing has the world's longest supply chains, energy networks have the world's highest node density, and financial institutions face layers of compliance requirements. These scenarios generate even larger data volumes, more diverse formats, and more uneven quality. For large models to truly enter these industries, they must bridge the gap from "general intelligence" to "industry intelligence." Underpinning this transition is precisely a solid, secure, and well-governed AI data infrastructure.
So how should this infrastructure be built? The practice of KeenData offers a case worth observing.
KeenData's core product is the KeenData Lakehouse, with a clear positioning: the industry's first integrated Data & AI intelligent-driven architecture.
What does "Data & AI integration" mean? It is not a traditional lakehouse with AI bolted on — that approach is like putting a new engine into an old car; it runs, but not well. It is an architecture designed natively for AI, called "AI-in-Lakehouse," where data is ready to be consumed by models from the moment it is stored.
Broken down, it does three things:
First, it connects the full chain. From data engineering → model training and inference → Agent factory → intelligent applications, it builds an end-to-end integrated technical closed loop covering data fusion, intelligent governance, model training and inference, and the full lifecycle of Agent operations. One platform connects every step from data to AI.
Second, it is powered by three core technologies. AI-native open lakehouse deeply integrates vector databases and multimodal storage capabilities, natively supporting hybrid queries. Multimodal alignment and fusion eliminate semantic gaps between text, images, audio, and video. Extreme training and inference performance is achieved through full-stack GPU acceleration — operator fusion, memory scheduling, mixed-precision training — squeezing every ounce of compute power out of the hardware.
Third, the product portfolio is comprehensive. It covers data integration, multimodal computing, data governance, AI model training, Agentic collaborative application development, scheduling, and deployment. With over 97% self-developed code, it meets localization and security compliance requirements, and is fully compatible with the domestic software and hardware ecosystem.
If you compare this with Databricks' architectural evolution, the consistency is striking.
At the 2026 DAIS Summit, Databricks' core direction was also to thicken the foundation and thin the Agent Runtime — LTAP enables OLTP and OLAP to share the same data, Lakehouse//RT delivers sub-100ms real-time queries, and Unity AI Gateway puts governance up front.
Two companies, separated by the Pacific Ocean, moving in the same direction with different technology stacks and product forms.
Databricks addresses the data infrastructure needs of global cloud-native scenarios; KeenData addresses the data infrastructure needs of large Chinese organizations in private deployment scenarios. The underlying technical logic is the same, but the implementation environments differ.
"Model-Data Resonance": A Two-Way Street — Data for AI + AI for Data
It is also worth noting that KeenData has a core methodology called "Model-Data Resonance."
In 2026, the Ministry of Industry and Information Technology (MIIT) and the National Data Bureau jointly launched the "Model-Data Resonance" initiative, and KeenData is one of its practitioners.
What does "Model-Data Resonance" mean? Broken down, it has two layers:
Data for AI — Industry data, after high-quality governance, feeds back into models, enabling large models to "understand industry jargon and know the rules." For example, historical operational data from petroleum refining helps process models understand industry boundaries; sensor data from automotive production lines enables vision models to detect micro-defects like seasoned experts.
AI for Data — Large models, in turn, activate dormant unstructured data. For instance, healthcare large models automatically extract disease-specific features from medical records to build high-quality clinical datasets; educational large models analyze teaching behaviors to generate personalized learning paths.
This is not a one-way "data-feeding-model" flow — it is a two-way street. Data becomes a flowing intelligent fuel, and the model becomes an alchemical furnace for data.
Moreover, this methodology has already proven itself with leading clients. Among KeenData's customer cases, Sinopec's data foundation has been recognized by the State-owned Assets Supervision and Administration Commission (SASAC) as a benchmark for digital transformation among central SOEs, with over 30 SOE representatives attending to hear the report. In addition, trusted data spaces and high-quality dataset solutions for large models and Agentic applications have been deployed at scale across multiple cities nationwide.
Coincidentally, Databricks also released Genie Ontology at the 2026 DAIS Summit — an automatically growing enterprise context graph, with the same core goal of making models "understand enterprise data." One calls it "Model-Data Resonance," the other calls it "contextual ontology." They are essentially talking about the same thing: models need to understand enterprise data, and data needs to be usable by models. The directions converge, the paths differ, but the underlying logic is identical.
A Trillion-Dollar Blue Ocean Market Has Just Opened
Now that we've covered who KeenData is, what it has done, and what it has proven, a larger question naturally emerges: how big is this market? What kind of track is this company entering?
The answer may exceed many people's expectations.
1. From "Application Procurement" to "Infrastructure Rebuilding"
Zoom out to the industry level.
Over the past decade, the typical scenario for enterprise data procurement was buying BI tools, buying reports, buying data warehouse tools. But today, customer needs have changed — with large models and Agents arriving, legacy platforms are no longer sufficient.
AI cannot directly consume raw data. Agents cannot bypass enterprise permissions and processes. Large model outputs must be supported by trustworthy data. What enterprises need is no longer one or two tools, but a complete system that can continuously supply high-quality data, continuously iterate model capabilities, and continuously produce intelligent applications.
In KeenData's words: "from selling software to delivering action systems, from building platforms to operating ecosystems."
So the AI data infrastructure market is not a red ocean like ERP or databases — it is a completely new blue ocean created by AI.
How big is the market? Extending from central SOEs down to their second- and third-tier subsidiaries, and from finance and energy to healthcare, transportation, and urban governance, the scale can be magnified dozens of times over.
2. Aligned with National Strategy
More importantly, this market is resonating with national strategy.
KeenData has mapped policy evolution into three phases: from "institutionalizing data as a production factor" to "national data infrastructure construction" to "AI+ action." Three leaps over six years — each phase generating rigid demand.
Behind policy keywords like trusted data spaces, high-quality datasets, and city brains lies the same direction: data infrastructure has shifted from an enterprise's optional choice to a necessity for industrial upgrading.
KeenData has already been selected as one of the first batch of pilot organizations for trusted data spaces by the National Data Bureau, and is building benchmark urban data infrastructure projects in Beijing, Hangzhou, Suzhou, and other cities.
3. What Does AI Data Infrastructure Mean for Industry?
If the above is still just a business-level judgment, the deeper question is: what does AI data infrastructure mean for industry?
The PC era had operating systems. The internet era had clouds. The foundational infrastructure of the AI era is the data foundation.
Without a solid data foundation, large models are castles in the air, and agents are mere paper talk. Whoever controls the high-quality data supply system controls the core competitiveness of the AI era.
KeenData offers a strategic formula: Data Resources × AI Data Infrastructure × High-Quality Datasets × Industry Models × Agent Networks × Scenario Operations = Sustainable, Monetizable Intelligent Capacity.
Behind this formula lies a judgment: what is truly scarce in the AI era, in addition to models and computing power, is the infrastructure that enables data to continuously become intelligence, intelligence to continuously enter business, and business to continuously generate returns.
Conclusion
If we stretch the evolution of digital civilization across a thirty-year horizon, three waves of infrastructure build-out become clearly visible:
The first was network infrastructure. Fiber optics and switches built the information superhighway, connecting the world.
The second was computing infrastructure. Cloud and data centers made computing power as accessible as electricity and water — enabling the world to "compute."
The third is now underway — data infrastructure. For the first time, it aims to turn the knowledge, experience, and production data accumulated by humanity over millennia into a "public cognitive layer" that AI can understand and call upon.
Whether it's Databricks, Snowflake, Palantir in the US, or KeenData in China — all of them have stepped precisely onto the starting point of this third infrastructure wave. What they are doing is laying the final missing piece for the large-scale commercialization of AI: an intermediate layer that enables data to be understood by models and models to be used by industry.
Models provide the brain. Compute provides the heartbeat. The data foundation provides the nervous system that makes everything truly work. Only when the three work in concert can AI truly leave the lab and enter every corner of factories, hospitals, energy sites, and urban governance.
This step will determine whether China can have its own roadbed in this underlying race. And it will determine whether the industrial value of AI is a passing craze, or a force that continues to change the world.
Back
Next article