A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
AI collaborators for science
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
*AI collaborators for science*
Summary This survey reviews over 270 datasets and 190 benchmarks for scientific LLMs (Sci-LLMs). It traces the evolution of models from data-focused pretraining to agentic frontiers where LLMs autonomously propose, test, and refine hypotheses
💡 Intuition Science is not just reading papers; it’s about experimenting and validating. Sci-LLMs should be more like lab partners than search engines.
🎯 Problem General LLMs are poorly equipped for multimodal, uncertainty-rich, and domain-specific data, and existing evaluations don’t capture scientific workflows.

🛠️ Solution


The survey organizes Sci-LLMs into:
- Data-Centric: pretrained on large, domain corpora.
- Domain-Specific Models: tailored for fields like chemistry or physics.
- Agentic Sci-LLMs: interactive models that design experiments and run analyses. It advocates for closed-loop ecosystems where LLMs don’t just process knowledge but also actively contribute to discovery. The paper also emphasizes reproducibility, open data, and benchmark standardization to fairly assess progress across disciplines.
🚧 Limitations and Future Opportunities
- Data scarcity and biases remain.
- Future: multimodal integration (text, images, simulation data) and autonomous “lab-bot” systems.
메타데이터
- post_id
- 7b19e1a6d5c8
- slug
- a-survey-of-scientific-large-language-models-from-data-foundations-to-agent-frontiers-7b19e1a6d5c8
- url
- https://medium.com/@huguosuo/a-survey-of-scientific-large-language-models-from-data-foundations-to-agent-frontiers-7b19e1a6d5c8
- canonical_url
- https://medium.com/@huguosuo/a-survey-of-scientific-large-language-models-from-data-foundations-to-agent-frontiers-7b19e1a6d5c8
- author_url
- https://medium.com/@huguosuo
- status
- ok
- fetched_at
- 2026-06-09 15:37:30