Parquet Schema Extractor for S3
Extracts and validates Parquet file schemas from Amazon S3 using the PyArrow library and AWS S3 SDK (boto3). Compares schemas across multiple partitions to detect schema drift and incompatible type changes. Outputs a schema diff report with partition paths and affected column details.
Parquet Schema Extractor for S3
Extracts and validates Parquet file schemas from Amazon S3 using the PyArrow library and AWS S3 SDK (boto3). Compares schemas across multiple partitions to detect schema drift and incompatible type changes. Outputs a schema diff report with partition paths and affected column details.
Installation
Method 1, Agent Skill Exchange
- Install from the marketplace listing: https://agentskillexchange.com/skills/parquet-schema-extractor-for-s3/
Method 2, Git clone
git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/parquet-schema-extractor-for-s3
Method 3, Download ZIP
- Download the repository ZIP and extract
skills/parquet-schema-extractor-for-s3.
Method 4, Manual copy
- Copy this skill folder into your local skills directory, then reload your agent tooling.
Method 5, Fork and sync
- Fork the repository if you want to maintain local edits while syncing upstream changes.