Trafilatura Web Text Extraction and Crawling Toolkit
Trafilatura is a Python package and CLI tool for gathering text from the web. It handles crawling, downloading, and extracting main text content, metadata, and comments from raw HTML, outputting clean structured data in CSV, JSON, Markdown, XML, and TXT formats.
Trafilatura Web Text Extraction and Crawling Toolkit
Trafilatura is a Python package and CLI tool for gathering text from the web. It handles crawling, downloading, and extracting main text content, metadata, and comments from raw HTML, outputting clean structured data in CSV, JSON, Markdown, XML, and TXT formats.
Installation
Method 1, Agent Skill Exchange
- Install from the marketplace listing: https://agentskillexchange.com/skills/trafilatura-web-text-extraction-crawling/
Method 2, Git clone
git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/trafilatura-web-text-extraction-crawling
Method 3, Download ZIP
- Download the repository ZIP and extract
skills/trafilatura-web-text-extraction-crawling.
Method 4, Manual copy
- Copy this skill folder into your local skills directory, then reload your agent tooling.
Method 5, Fork and sync
- Fork the repository if you want to maintain local edits while syncing upstream changes.