What it does

  • Extracts title, canonical URL, language, and useful meta tags
  • Preserves headings, paragraphs, lists, links, tables, and code blocks
  • Captures figures, media, forms, and definition lists
  • Parses JSON-LD structured data when present
  • Lists external stylesheets and scripts
  • Removes hidden elements, navigation clutter, and common boilerplate
  • Keeps Facebook noise filtering without being tied to Facebook-only HTML
  • Outputs cleaner, more complete, human-readable text