Extract schema.org, Open Graph, and JSON-LD metadata from web pages for indexing

Uses extruct to pull machine-readable metadata from raw HTML so an agent can classify, deduplicate, or enrich pages without brittle full-page parsing. It is best for metadata harvesting workflows, not for crawling an entire site or rendering JavaScript-heavy pages.

Extract schema.org, Open Graph, and JSON-LD metadata from web pages for indexing

Uses extruct to pull machine-readable metadata from raw HTML so an agent can classify, deduplicate, or enrich pages without brittle full-page parsing. It is best for metadata harvesting workflows, not for crawling an entire site or rendering JavaScript-heavy pages.

Installation

Method 1, Agent Skill Exchange

Method 2, Git clone

git clone https://github.com/agentskillexchange/skills.git && cd skills/skills/extract-schema-org-open-graph-and-json-ld-metadata-from-web-pages-for-indexing

Method 3, Download ZIP

  • Download the repository ZIP and extract skills/extract-schema-org-open-graph-and-json-ld-metadata-from-web-pages-for-indexing.

Method 4, Manual copy

  • Copy this skill folder into your local skills directory, then reload your agent tooling.

Method 5, Fork and sync

  • Fork the repository if you want to maintain local edits while syncing upstream changes.

Source