A fast HTML5 parser with CSS selectors, written in Cython, using the Lexbor engine.
Installation
From PyPI using pip:
pip install selectolax
If installation fails due to compilation errors, you may need to install Cython:
pip install selectolax[cython]
This usually happens when you try to install an outdated version of selectolax on a newer version of Python.
Development version from GitHub:
git clone --recursive https://github.com/rushter/selectolax
cd selectolax
pip install -r requirements_dev.txt
python setup.py install
How to compile selectolax while developing:
make clean
make dev
Basic examples
Here are some basic examples to get you started with selectolax:
Parsing HTML and extracting text:
from selectolax.lexbor import LexborHTMLParser
html = """
<h1 id="title" data-updated="20201101">Hi there</h1>
<div class="post">Lorem Ipsum is simply dummy text of the printing and typesetting industry. </div>
<div class="post">Lorem ipsum dolor sit amet, consectetur adipiscing elit.</div>
"""
tree = LexborHTMLParser(html)
print(tree.css_first('h1#title').text())
'Hi there'
print(tree.css_first('h1#title').attributes)
{'id': 'title', 'data-updated': '20201101'}
print([node.text() for node in tree.css('.post')])
['Lorem Ipsum is simply dummy text of the printing and typesetting industry. ',
'Lorem ipsum dolor sit amet, consectetur adipiscing elit.']
Using advanced CSS selectors
from selectolax.lexbor import LexborHTMLParser
html = "<div><p id=p1><p id=p2><p id=p3><a>link</a><p id=p4><p id=p5>text<p id=p6></div>"
selector = "div > :nth-child(2n+1):not(:has(a))"
for node in LexborHTMLParser(html).css(selector):
print(node.attributes, node.text(), node.tag)
print(node.parent.tag)
print(node.html)
{'id': 'p1'} p
div
<p id="p1"></p>
{'id': 'p5'} text p
div
<p id="p5">text</p>
Using lexbor-contains CSS pseudo-class to match text
from selectolax.lexbor import LexborHTMLParser
html = "<div><p>hello </p><p id='main'>lexbor is AwesOme</p></div>"
parser = LexborHTMLParser(html)
Case-insensitive search
results = parser.css('p:lexbor-contains("awesome" i)')
Case-sensitive search
results = parser.css('p:lexbor-contains("AwesOme")')
assert len(results) == 1
assert results[0].text() == "lexbor is AwesOme"
Simple Benchmark
- Extract title, links, scripts and a meta tag from main pages of top 754 domains. See
examples/benchmark.pyfor more information.
Links
- selectolax API reference and examples
- Video introduction to web scraping using selectolax
- How to Scrape 7k Products with Python using selectolax and httpx
- Lexbor benchmark
- Python benchmark
- Another Python benchmark
- Universal interface to lxml and selectolax
License
- lexbor engine — Apache-2.0 license
- selectolax - MIT
Contributors
Thanks to all the contributors of selectolax!