← Back to list

The New Oil: Data in the AI Revolution

The New Oil: Data in the AI Revolution

JUJALU · 2025-02-25 11:42 · 0 claps · 2.2 min read
#machine-learning #ai #ai-industries #ai-robotics #ai-biology
Open on Medium ↗
Wiki topics: ML · Machine Learning AI · AI · General EDU · Education & Learning

The New Oil: Data in the AI Revolution

The New Oil: Data in the AI Revolution

After spending considerable time in the Bay Area immersing myself in the AI ecosystem, I’ve gained clarity on the direction of AI development. One thing stands clear: data is indeed the new oil of the AI industrial revolution. While the paths to acquiring this crucial resource traditionally involve public data collection and web scraping, the future might lie in something more specific: vertical models.

The Rise of Vertical Models

Recent conversations, including insights from Gensen’s interview, point to the growing importance of vertical models — AI systems specialized for specific domains. These specialized models are emerging in several key areas:

Robotics and Vision

Several promising developments are worth noting:

  • Liquid neural networks for real-life learning
  • Continuous video updates for better understanding
  • Key players in the humanoid robotics space:
  • k-scale labs
  • Red Rabbit Robotics (open-source humanoid)
  • TAU Robotics
  • The Jetson Orin Nano developer kit as an enabling platform

Biology and Life Sciences

Perhaps one of the most exciting vertical spaces is biological AI, where companies like Profluent are developing specialized language models for biology. This space showcases the difference between general and specialized AI architectures:

Comparing Architectures:

AlphaFold:

  • Specialized for protein structure prediction
  • Uses Evoformer blocks
  • Processes amino acid sequences and alignments
  • Produces 3D coordinate predictions
  • Incorporates physical and biological constraints

Traditional LLMs (like LLaMA):

  • General-purpose language processing
  • Based on Transformer architecture
  • Processes text sequences
  • Generates textual outputs
  • Relies on statistical patterns

The Strategy for Vertical Data Collection

The path to building effective vertical models follows a clear pattern:

  1. Gather initial minimal data to create a baseline product
  2. Launch to market when minimally viable
  3. Collect more real-world data through deployment
  4. Iterate and improve based on actual usage

Practical Approaches to Model Development

The most practical approach to training large AI models mirrors human learning: building upon existing knowledge. Just as you would teach a 15-year-old based on their current understanding, fine-tuning existing large language models is often more efficient than starting from scratch.

The Biology Model Example

In biological applications, we’re seeing interesting developments in:

  • Cell morphology analysis
  • Longitudinal studies
  • Transcriptomic data processing
  • Environment-dependent cell state prediction
  • Integration with human IPS cells

The future holds immense potential for AI in biology:

  • Automated literature review and understanding
  • AI-driven experiment design and execution
  • Automated image inspection and analysis
  • Statistical analysis automation
  • Clinical trial prediction

Looking Forward

The convergence of specialized architectures, domain expertise, and targeted data collection is creating new opportunities in AI development. Whether it’s robotics in Asia or biological AI in the US, the field is rapidly evolving toward more specialized, efficient solutions.

The key to success in this new landscape isn’t just about having the biggest model or the most data — it’s about having the right data and the right architecture for your specific domain. As we move forward, the ability to effectively combine domain expertise with AI capabilities will become increasingly valuable.

Conclusion

My time in the Bay Area has shown me that while general-purpose AI models will continue to be important, the real breakthroughs are likely to come from specialized vertical applications. Whether joining an existing startup or creating something new, understanding this landscape of specialized AI is crucial for making informed decisions about future projects.

This article reflects personal observations from recent immersion in the Bay Area AI ecosystem and conversations with industry leaders.


메타데이터
post_id
2c8a36b920bd
slug
the-new-oil-data-in-the-ai-revolution-2c8a36b920bd
url
https://medium.com/@JUJALU/the-new-oil-data-in-the-ai-revolution-2c8a36b920bd
canonical_url
https://medium.com/@JUJALU/the-new-oil-data-in-the-ai-revolution-2c8a36b920bd
author_url
https://medium.com/@JUJALU
status
ok
fetched_at
2026-06-09 15:37:30