Lightweight parquet reference files that turn 3--4 GB ECMWF GRIB files into ~140 KB virtual datasets, enabling Dask-based parallel analysis without downloading the raw data.
Each parquet file contains a table of [zarr_key, [s3_url, byte_offset, byte_length]] references pointing into ECMWF IFS ensemble GRIB files on AWS S3 (s3://ecmwf-forecasts/). Instead of… See the full description on the dataset page:
https://huggingface.co/datasets/E4DRR/gik-ecmwf-par.