22 AI for Science Datasets Spanning Genomics, Materials, Climate & Agriculture

This article compiles 22 AI for Science datasets across biomedical, agricultural, materials chemistry, and earth science domains, detailing each dataset's source, size, and application for model training and benchmarking.

HyperAI Super Neural
HyperAI Super Neural
HyperAI Super Neural
22 AI for Science Datasets Spanning Genomics, Materials, Climate & Agriculture

AI for Science is expanding how artificial intelligence participates in scientific research. From gene analysis and crop disease identification to material property prediction and meteorological modeling, AI is gradually integrating into research workflows across disciplines, helping researchers identify patterns, propose hypotheses, and conduct validation from complex data.

As AI and domain sciences converge, the value of scientific data extends to model training, capability evaluation, and experimental optimization. For example, LLM4Mat-Bench provides benchmark resources for material property prediction, while Reac-Discovery combines structural design, experimental data, and machine learning to optimize reactors and reaction conditions. Data is connecting scientific problems, computational models, and experimental verification, enabling AI to participate more deeply in the research process.

This article organizes 22 AI for Science datasets into four categories: biomedical, agricultural science, materials chemistry, and earth science. They cover gene case analysis, disease image recognition, crystal structures and material properties, meteorological and hydrological observations, and ecological environment research. These resources provide entry points for data analysis, model training, and benchmarking across fields, highlighting the need for continuous construction of data foundations with clear provenance, reliable quality, and suitability for specific research tasks.

Biomedical

1. TACK – Targeting Chimeras Knowledge Dataset

Online: https://go.hyper.ai/jbIBf TACK (TArgeting Chimeras Knowledge) is a standardized knowledge base and benchmark dataset released in 2026 by the AI Laboratory for Molecular Engineering. It is built for machine learning-driven PROTAC degradation activity prediction, addressing data scarcity, lack of rigorous evaluation, and limited coverage in existing PROTAC ML benchmarks. The dataset contains 6,561 records covering 3,514 unique PROTAC molecules, 164 target proteins (POI), 9 E3 ubiquitin ligases, and 155 cell lines, offering rich chemical structural features and biological experimental condition diversity.

2. GeneBench-Pro Public Package – Gene Case Benchmark Dataset

Online: https://go.hyper.ai/CooDU Released by OpenAI, this public gene analysis benchmark dataset provides a reproducible research case study evaluation environment for AI agents in genomics and biomedicine. It includes 10 independent problem cases spanning statistical genetics, clinical genomics, population genetics, single-cell analysis, 3D genomics, and functional genomics. Each case contains task descriptions, public reference answers (ground truth), scoring rules (grader contract), and an official technical report.

3. Sleep – Mammalian Sleep Traits Dataset

Online: https://go.hyper.ai/FTmEG Sleep is a mammalian sleep trait dataset for statistical analysis in sleep ecology and comparative physiology, associated with the paper "Sleep in Mammals: Ecological and Constitutional Correlates." It includes brain weight, body weight, lifespan, gestation period, sleep duration, predation and danger indices for 62 mammalian species.

4. SASH-VPV – Subcutaneous Palm Vein Recognition Dataset

Online: https://go.hyper.ai/d8lfc SASH-VPV is a near-infrared palm vein biometric benchmark dataset for biometric identification and computer vision research, focusing on identity authentication using subcutaneous vein structures. It is widely used in biometric system development, deep learning model training, and cross-session robustness studies.

SASH-VPV dataset example
SASH-VPV dataset example

Dataset example

5. Eczema & Tinea Skin Disease – Eczema Skin Disease Dataset

Online: https://go.hyper.ai/jeBG8 This medical image dataset of eczema and tinea skin diseases provides concise, practical data for binary image classification tasks. It contains 2,147 skin disease images and is used for skin disease image classification, deep learning model training and evaluation, few-shot and transfer learning research, and medical image analysis teaching.

Agricultural Science

1. Rose Leaf Disease – Rose Leaf Disease Dataset

Online: https://go.hyper.ai/QhBW1 Rose Leaf Diseases provides high-quality rose leaf images for developing and benchmarking rose leaf disease detection models, supporting plant monitoring system construction. The original version contains 2,458 images from Bangladesh, categorized into five classes: black spot, downy mildew, dry leaf, healthy, and insect hole.

2. Rice Leaf Diseases – Rice Leaf Disease Detection Dataset

Online: https://go.hyper.ai/rS1FT Built for precision agriculture object detection, this dataset enhances computer vision models for real-world agricultural disease detection, classification, and localization. It is widely used for YOLO series training, agricultural disease detection, edge vision deployment, and intelligent rice planting management. It contains 8,665 rice leaf images across 9 categories: healthy leaves and 8 common diseases (bacterial leaf blight, brown spot, rice leaf folder, rice blast, leaf scald, false smut, narrow brown spot, neck blast).

3. GRAPE Leaf Diseases – Grape Leaf Disease Detection Dataset

Online: https://go.hyper.ai/PsaWZ Designed for precision agriculture object detection, this dataset improves grape leaf disease detection, classification, and localization in real agricultural scenarios. It contains 4,195 grape leaf images across 4 categories: healthy leaves and 3 common diseases (black rot, esca, leaf blight).

4. Corn Leaf Diseases – Corn Leaf Disease Detection Dataset

Online: https://go.hyper.ai/WpVFZ Designed for precision agriculture object detection, this dataset enhances disease detection and classification in agricultural scenes. It contains 4,027 corn leaf images across 4 categories: healthy leaves and 3 common diseases (rust, gray spot, wilt).

Corn leaf disease dataset example
Corn leaf disease dataset example

Dataset example

5. Apple Leaf Diseases – Apple Leaf Disease Detection Dataset

Online: https://go.hyper.ai/rivXt This high-quality apple leaf image dataset is designed for precision agriculture object detection and can be directly used for YOLOv8, YOLOv11 training, agricultural disease recognition research, and smart agriculture application development. It contains 3,444 apple leaf images across 4 categories: healthy leaves and 3 high-incidence diseases (black rot, cedar rust, scab).

Apple leaf disease dataset example
Apple leaf disease dataset example

Dataset example

Materials Chemistry

1. Reac-Discovery – Chemical Reactor Performance Dataset

Online: https://go.hyper.ai/STqAK Released in 2025 by Jaume I University, this dataset supports AI-driven flow reactor design and reaction performance optimization. The associated paper is "Reac-Discovery: an artificial intelligence–driven platform for continuous-flow catalytic reactor discovery and optimization." Data was automatically generated by the team's Reac-Discovery platform without external public sources. It covers geometric structure, printability, and reaction performance data, corresponding to the platform's Reac-Gen, Reac-Fab, and Reac-Eval modules.

2. QMOF150 – Quantum Chemistry Dataset

Online: https://go.hyper.ai/2e3WA Released in 2025 by Meta and Cambridge University, this quantum chemistry dataset aims to accelerate quantum material discovery. The associated paper is "All-atom Diffusion Transformers: Unified generative modelling of molecules and materials." It contains ~14,000 metal-organic frameworks (MOFs) and coordination polymers. Experimentally characterized MOFs, after DFT structural relaxation, include computed properties such as optimized geometry, energy, band gap, charge density, density of states, partial charges, spin density, and bond order.

3. MP-20-PXRD – Atomic Material Benchmark Dataset

Online: https://go.hyper.ai/XndaJ Proposed in 2025 by Columbia University and Stanford University for end-to-end training of PXRDnet, a diffusion-model-based generative AI structure solution method. The research, "Ab initio structure solutions from nanocrystalline powder diffraction data via diffusion models," was published in Nature Materials. The dataset samples material compositions from the Materials Project database with up to 20 atoms per unit cell, comprising 45,229 materials split 90%/7.5%/2.5% for training, validation, and testing.

4. LLM4Mat-Bench – Crystal Structure Dataset

Online: https://go.hyper.ai/uHA6X Created jointly by Princeton University, University of Toronto, and others, this multimodal language model evaluation dataset benchmarks large language models (LLMs) for material property prediction and discovery. The paper "LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction" describes it. It collects ~1.97 million crystal structure samples from 10 public material databases, covering 45 distinct material physical and chemical properties, making it the largest benchmark for evaluating LLMs on material property prediction.

LLM4Mat-Bench statistics
LLM4Mat-Bench statistics

LLM4Mat-Bench statistics

5. Material DFT – Material Property Dataset

Online: https://go.hyper.ai/LbIWK A curated dataset for materials science R&D, providing high-quality material property records from the Materials Project database. It covers diverse chemical compositions and physical properties, each corresponding to a unique material. All properties are computed via Density Functional Theory (DFT), a widely used computational method for predicting material behavior. Suitable for material property modeling, machine learning training, and material discovery.

Earth Science

1. Eclipse Soundscapes ESID #657 – Total Solar Eclipse Soundscape Audio Dataset

Online: https://go.hyper.ai/csQWG Released by the NASA-funded citizen science project Eclipse Soundscapes, this 2024 total solar eclipse soundscape audio dataset provides raw recordings for studying environmental acoustic changes and animal vocal behavior during eclipses. It supports ecoacoustic analysis, eclipse ecological effect research, and citizen science data quality studies. Audio was recorded at station ESID #657 (latitude 37.48431°, longitude -89.089°) on April 8, 2024, and surrounding days (April 6–11, 2024, UTC).

2. USGS Global Earthquake – Global Regional Earthquake Dataset

Online: https://go.hyper.ai/dWn7v Compiled from USGS earthquake event data, this dataset supports spatiotemporal distribution analysis of seismic activity, regional seismic hazard research, and data visualization teaching. It covers 1900–2026 earthquake records. Each file records a single event with magnitude, origin time, latitude, longitude, depth, epicenter description, and other standard USGS fields, supplemented with continent, region, and country fields based on coordinates. The dataset is a regionally split version of the raw global earthquake data, saved as independent CSV files per geographic region.

USGS earthquake dataset example
USGS earthquake dataset example

Dataset example

3. Maritime Continent Satellite Weather – Maritime Continent Satellite Meteorological Dataset

Online: https://go.hyper.ai/T18Kp This satellite meteorological image dataset for the Maritime Continent region explores deep learning applications in weather forecasting and nowcasting, supporting rainfall prediction, thunderstorm identification, and multimodal model fusion. It contains 95,036 hourly Himawari-8/9 satellite infrared image frames paired with meteorological data, covering the western Maritime Continent (Southeast Asia and Malacca Strait, including Singapore, Indonesia, Malaysia, Thailand). It includes ground weather data for 23 cities in these countries, each with 97,056 rows.

4. Global Climate & Energy Transition 2000-2026 – Global Climate and Energy Transition Dataset

Online: https://go.hyper.ai/Vjtpq This dataset systematically characterizes global climate change and energy transition processes for climate science, ESG research, energy economics, carbon market analysis, policy evaluation, and machine learning prediction. It records 2000–2026 global climate and energy transition, including global and regional temperature anomalies, CO2 emissions for 50 countries, energy structure transition (coal to renewables), prices of 5 major carbon emission trading markets, and 50 representative major climate events.

5. Caravan – Global Community Large-Sample Hydrological Dataset

Online: https://go.hyper.ai/e3nmq Caravan is an open global community large-sample hydrological dataset that standardizes and integrates seven existing large-sample hydrological datasets. It contains meteorological forcing data, streamflow data, and static catchment attributes (geophysical, societal, climatic) for 6,830 catchments globally. Data sources include CAMELS (US), CAMELS-AUS, CAMELS-BR, and others. Streamflow data is normalized by catchment area (mm/day). All data are recorded in local time zones (no daylight saving), with time zone information in metadata.

6. Aquatic Wildlife Atlas – Global Aquatic Species Records Dataset

Online: https://go.hyper.ai/VOfsw This large-scale aquatic animal observation dataset supports aquatic ecological research and biodiversity analysis. It provides 200,000 aquatic animal observation records covering 100+ aquatic species across major global aquatic environments: coral reefs, tropical rivers, Arctic waters, and deep-sea regions down to 7,000 meters. Each record represents an independent biological observation with complete Linnaean taxonomy (kingdom to species), ecological traits, physical measurements, geographic coordinates, environmental readings, and observation metadata.

7. Global Earthquake-M4.5 – Global M4.5+ Earthquake Dataset

Online: https://go.hyper.ai/vwLoS Built for seismic activity analysis and geospatial research, this dataset helps analyze long-term seismic frequency, distribution, and magnitude changes. It is used in seismic hazard assessment, urban planning, geological research, and disaster early warning. It contains 230,608 earthquake records covering global M4.5+ events from 1900 to 2026.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

machine learningdatasetsbenchmarkgenomicsAI for Sciencematerials scienceagricultureclimate science
HyperAI Super Neural
Written by

HyperAI Super Neural

Deconstructing the sophistication and universality of technology, covering cutting-edge AI for Science case studies.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.