How to Extract Text from HTML with Python: 4 Reliable Methods
TL;DR
Beautiful Soup is the clearest general-purpose method for extracting text from imperfect HTML.
lxml is a strong choice when XPath and high-throughput parsing matter.
Trafilatura is better when the goal is main article text rather than every visible navigation label.
Inscriptis is useful when text layout, tables, and lists need better preservation.
Remove scripts and style content, preserve block boundaries, normalize Unicode, and validate against expected text before storage.
What is the best way to extract text from HTML with Python?
The best method depends on whether the input is a snippet, a complete document, or a rendered page. Nstdata Crawl can provide cleaned or rendered source content when acquisition is the hard part; Beautiful Soup, lxml, Trafilatura, and Inscriptis solve different parsing problems after HTML is available. Regex alone is not a reliable HTML parser.
The Python web-scraping project guide provides broader pipeline context. Text extraction should preserve enough structure for its destination: search indexing may need headings and paragraphs, while a compact classifier may need normalized plain text.
What do you need before you start?
Install the four libraries in an isolated environment. Use one shared fixture so the outputs can be compared fairly.
Method 1: Use Beautiful Soup for readable general-purpose parsing
Step 1: Build the parse tree
from bs4 import BeautifulSoup
soup = BeautifulSoup(HTML,"html.parser")
Step 2: Remove non-content nodes
for node in soup.select("script, style, noscript, template"): node.decompose()
Step 3: Preserve boundaries and normalize whitespace
lines =[line.strip()for line in soup.get_text("\n").splitlines()]text ="\n".join(line for line in lines if line)print(text)
Beautiful Soup tolerates malformed markup and has an approachable selector API. Its limitation is that get_text() does not know which visible blocks are primary content. Use semantic selectors such as main, article, or a verified content container when available. Check the official Beautiful Soup documentation for parser behavior.
Method 2: Use lxml for XPath and throughput
Step 1: Parse the document
from lxml import html
tree = html.fromstring(HTML)
Step 2: Remove unwanted elements
for node in tree.xpath("//script|//style|//noscript|//template"): node.drop_tree()
Step 3: Extract selected text nodes
parts = tree.xpath("//main//text()[normalize-space()]")text ="\n".join(part.strip()for part in parts if part.strip())
lxml is fast and gives precise XPath control, but careless XPath can flatten meaningful structure or collect hidden content. The official lxml.html documentation describes its HTML helpers.
Method 3: Use Trafilatura for main-content extraction
Step 1: Pass the complete document
import trafilatura
text = trafilatura.extract(HTML, include_links=False, include_images=False)
Step 2: Handle a missing result
Trafilatura can return no main content for short, unusual, or navigation-heavy pages. Treat that as a validation failure and fall back to a scoped DOM method rather than storing an empty document.
Step 3: Keep metadata separately
Main-content extraction deliberately removes boilerplate. If title, canonical URL, author, or publication date matters, extract those fields into metadata instead of expecting them to survive plain-text conversion. Verify current options in the Trafilatura documentation.
Method 4: Use Inscriptis when layout carries meaning
Step 1: Convert HTML to layout-aware text
from inscriptis import get_text
text = get_text(HTML)
Step 2: Review lists and tables
Inscriptis aims to preserve visible layout better than a simple descendant-text join. This can help with tables, lists, and documents where line breaks communicate structure.
Step 3: Normalize for the destination
Do not apply aggressive whitespace collapsing if table columns or list indentation matter. Create separate normalization profiles for search, LLM ingestion, and archival review.
How do you fetch HTML before extracting text?
Use a bounded HTTP client for static pages and validate status, content type, encoding, final URL, and expected page identity. If the page requires JavaScript, use a browser or managed rendering layer before parsing. Nstdata Crawl is relevant when rendering and source acquisition should be managed; consult current Crawl pricing before estimating production cost.
The JavaScript web-scraping guide explains when an HTML parser cannot see browser-generated content. A parser cannot recover data that never arrived in its input.
For larger jobs, the web-scraping scaling guide explains why parser throughput must be evaluated together with queues, retries, and accepted-record quality.
How do you clean extracted text without damaging it?
Normalize Unicode, convert non-breaking spaces, remove repeated blank lines, and retain block boundaries. Avoid deleting all short lines because headings and labels can be short. Avoid lowercasing content intended for display or entity extraction. Keep the raw source or a permitted hash so a cleaning regression can be investigated.
Build golden fixtures containing headings, paragraphs, entities, lists, tables, hidden nodes, malformed tags, and non-ASCII text. Compare exact output or a structured block representation in automated tests. The WHATWG HTML standard is the primary reference when parser behavior depends on document structure.
Conclusion
Use Beautiful Soup for understandable general parsing, lxml for XPath and performance, Trafilatura for article-body extraction, and Inscriptis when layout matters. Evaluate methods against the same fixtures and measure missing content, boilerplate, and structural preservation. When the input HTML itself is incomplete, solve acquisition first with a browser or Nstdata Crawl. For repeated jobs with multiple proxy sources, Nstdata Proxy Manager can address routing separately from parsing.
Experience Nstdata β Start Your Free Trial Today
Yes. Beautiful Soup and lxml both recover many malformed documents, but their repairs may differ, so test the selected parser against representative fixtures.
Q: Why does get_text include JavaScript or CSS?
The parser treats script and style contents as text nodes unless those elements are removed. Decompose them before extracting text.
Q: Is regex suitable for removing HTML tags?
Regex can clean a known fragment, but it is not a reliable general HTML parser because nesting, malformed markup, entities, and embedded languages complicate the grammar.
Q: How do you extract only article text?
Use a verified article or content selector when the template is known, or a main-content library such as Trafilatura when templates vary.
Q: Which method is fastest?
lxml and other native-backed parsers are often faster than Beautiful Soup, but throughput should be measured with the real documents and required cleanup.
Ivy Lin
Sep. 28th 2026
Crawl entire websites with a single API request
99.8% success rate with JavaScript rendering
Get clean, LLM-ready data in multiple formats
Turn any website into Markdown, HTML, JSON, links, PDFs and more β without managing crawling infrastructure.