discovered 03 Aug 2026
croissant
→ View on GitHubCroissant is a high-level format designed to standardize machine learning datasets by encapsulating metadata, resource file descriptions, data structures, and relevant ML semantics into a single file. It enhances dataset accessibility and usability through a structured approach based on schema.org's Dataset vocabulary. Notable features include easy integration with existing datasets, metadata-driven dataset management, and support for various ML tools, streamlining the dataset usage in machine learning workflows.