Master XPath Syntax for Web Scraping: Node Relationships & Expression Guide
This tutorial introduces XPath syntax for web scraping, explaining HTML node relationships (parent, child, sibling, ancestor, descendant) and demonstrating key expressions including path operators (/ vs //), attribute selection (@), predicates, wildcards, and union operators with practical examples for element targeting.
XPath enables locating specific elements in HTML or XML structures using a path-like syntax similar to Windows file paths, and includes a standard function library for powerful queries.
HTML Node Relationships
HTML documents have a hierarchical structure with defined relationships:
Parent node : The direct container of a node (e.g., <body> is parent of <nav>)
Child node : Direct descendant (e.g., <nav> is child of <body>)
Sibling nodes : Nodes at the same level (e.g., <nav> and first <div> are siblings; <li> tags on lines 177–181 are siblings)
Ancestor nodes : All nodes above in the hierarchy (parent, grandparent, etc.)
Descendant nodes : All nodes below in the hierarchy (children, grandchildren, etc.)
Core XPath Expressions
//@class— Select all attributes named
class /article— Select root element
article //div— Select all div child elements article — Select all child nodes of all article elements article/a — Select all a elements that are children of
article article//div— Select all div descendants of article Key distinction: / selects direct children only; // selects all descendants (broader scope). The @ prefix selects attributes (commonly @class).
Attribute-Based and Positional Selection
//div[@lang]— Select all div elements that have a lang attribute //div[@lang='eng'] — Select all div elements with lang attribute equal to
eng /article/div[1]— Select the first div child of
article /article/div[last()]— Select the last div child of
article /div/*— Select all child nodes of div elements //* — Select all elements in the document //div/a | //div/p — Select all a and p elements that are children of div (union operator)
In the HTML source, class is an attribute name and the value after = (e.g., grid-5) is the attribute value; a node can have multiple attributes.
Practical Application
After mastering these XPath expressions, you can write queries to extract target data from web pages. The next article will cover using XPath in Scrapy spider projects.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Python Crawling & Data Mining
Life's short, I code in Python. This channel shares Python web crawling, data mining, analysis, processing, visualization, automated testing, DevOps, big data, AI, cloud computing, machine learning tools, resources, news, technical articles, tutorial videos and learning materials. Join us!
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
