Turn messy document collections into structured rows with DocETL

Define repeatable extraction pipelines that pull fields from large document collections, normalize outputs, and audit failures across the corpus.

Turn messy document collections into structured rows with DocETL

Define repeatable extraction pipelines that pull fields from large document collections, normalize outputs, and audit failures across the corpus.

Installation

Method 1, Agent Skill Exchange

Method 2, Git clone

git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/turn-messy-document-collections-into-structured-rows-with-docetl

Method 3, Download ZIP

  • Download the repository ZIP and extract skills/turn-messy-document-collections-into-structured-rows-with-docetl.

Method 4, Manual copy

  • Copy this skill folder into your local skills directory, then reload your agent tooling.

Method 5, Fork and sync

  • Fork the repository if you want to maintain local edits while syncing upstream changes.

Source