Profile
Back to NewsBack
GitHub Trending 3 min
Reader Mode
rushter/selectolax: Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.

rushter/selectolax: Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.

9 hours ago

!selectolax logo


A fast HTML5 parser with CSS selectors, written in Cython, using the Lexbor engine.


PyPI - Version</a> PyPI Total Downloads</a> CI</a> GitHub License</a>


Installation

From PyPI using pip:

pip install selectolax

If installation fails due to compilation errors, you may need to install Cython:

pip install selectolax[cython]

This usually happens when you try to install an outdated version of selectolax on a newer version of Python.

Development version from GitHub:

git clone --recursive  https://github.com/rushter/selectolax
cd selectolax
pip install -r requirements_dev.txt
python setup.py install

How to compile selectolax while developing:

make clean
make dev

Basic examples

Here are some basic examples to get you started with selectolax:

Parsing HTML and extracting text:

from selectolax.lexbor import LexborHTMLParser

html = """ <h1 id="title" data-updated="20201101">Hi there</h1> <div class="post">Lorem Ipsum is simply dummy text of the printing and typesetting industry. </div> <div class="post">Lorem ipsum dolor sit amet, consectetur adipiscing elit.</div> """ tree = LexborHTMLParser(html)

print(tree.css_first('h1#title').text())

'Hi there'

print(tree.css_first('h1#title').attributes)

{'id': 'title', 'data-updated': '20201101'}

print([node.text() for node in tree.css('.post')])

['Lorem Ipsum is simply dummy text of the printing and typesetting industry. ',

'Lorem ipsum dolor sit amet, consectetur adipiscing elit.']

Using advanced CSS selectors

from selectolax.lexbor import LexborHTMLParser

html = "<div><p id=p1><p id=p2><p id=p3><a>link</a><p id=p4><p id=p5>text<p id=p6></div>" selector = "div > :nth-child(2n+1):not(:has(a))"

for node in LexborHTMLParser(html).css(selector): print(node.attributes, node.text(), node.tag) print(node.parent.tag) print(node.html)

{'id': 'p1'} p

div

<p id="p1"></p>

{'id': 'p5'} text p

div

<p id="p5">text</p>

Using lexbor-contains CSS pseudo-class to match text

from selectolax.lexbor import LexborHTMLParser
html = "<div><p>hello </p><p id='main'>lexbor is AwesOme</p></div>"
parser = LexborHTMLParser(html)

Case-insensitive search

results = parser.css('p:lexbor-contains("awesome" i)')

Case-sensitive search

results = parser.css('p:lexbor-contains("AwesOme")') assert len(results) == 1 assert results[0].text() == "lexbor is AwesOme"

Simple Benchmark

  • Extract title, links, scripts and a meta tag from main pages of top 754 domains. See examples/benchmark.py for more information.
| Package | Time | |-------------------------------|-----------| | Beautiful Soup (html.parser) | 61.02 sec.| | lxml / Beautiful Soup (lxml) | 9.09 sec. | | html5_parser | 16.10 sec.| | selectolax (Lexbor) | 2.39 sec. |

Links

License

Contributors

Thanks to all the contributors of selectolax!

Chat with me