Apache Tika Content Extraction Hub
Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.
Apache Tika Content Extraction Hub
Extracts text and metadata from 1400+ file formats via Apache Tika Server REST API. Handles PDF, DOCX, PPTX, email archives, and embedded document extraction with MIME type detection.
Installation
Method 1, Agent Skill Exchange
- Install from the marketplace listing: https://agentskillexchange.com/skills/apache-tika-content-extraction-hub/
Method 2, Git clone
git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/apache-tika-content-extraction-hub
Method 3, Download ZIP
- Download the repository ZIP and extract
skills/apache-tika-content-extraction-hub.
Method 4, Manual copy
- Copy this skill folder into your local skills directory, then reload your agent tooling.
Method 5, Fork and sync
- Fork the repository if you want to maintain local edits while syncing upstream changes.