Apache Tika Document Parser
Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.
Apache Tika Document Parser
Extracts structured text, metadata, and embedded objects from PDFs, Office documents, and 1000+ file formats using the Apache Tika REST API. Outputs clean Markdown or JSON with XMP metadata preservation.
Installation
Method 1, Agent Skill Exchange
- Install from the marketplace listing: https://agentskillexchange.com/skills/apache-tika-document-parser/
Method 2, Git clone
git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/apache-tika-document-parser
Method 3, Download ZIP
- Download the repository ZIP and extract
skills/apache-tika-document-parser.
Method 4, Manual copy
- Copy this skill folder into your local skills directory, then reload your agent tooling.
Method 5, Fork and sync
- Fork the repository if you want to maintain local edits while syncing upstream changes.