Master XPath Syntax for Web Scraping: Node Relationships & Expression Guide

This tutorial introduces XPath syntax for web scraping, explaining HTML node relationships (parent, child, sibling, ancestor, descendant) and demonstrating key expressions including path operators (/ vs //), attribute selection (@), predicates, wildcards, and union operators with practical examples for element targeting.

Python Crawling & Data Mining
Python Crawling & Data Mining
Python Crawling & Data Mining
Master XPath Syntax for Web Scraping: Node Relationships & Expression Guide

XPath enables locating specific elements in HTML or XML structures using a path-like syntax similar to Windows file paths, and includes a standard function library for powerful queries.

HTML Node Relationships

HTML documents have a hierarchical structure with defined relationships:

Parent node : The direct container of a node (e.g., <body> is parent of <nav>)

Child node : Direct descendant (e.g., <nav> is child of <body>)

Sibling nodes : Nodes at the same level (e.g., <nav> and first <div> are siblings; <li> tags on lines 177–181 are siblings)

Ancestor nodes : All nodes above in the hierarchy (parent, grandparent, etc.)

Descendant nodes : All nodes below in the hierarchy (children, grandchildren, etc.)

Node relationship diagram
Node relationship diagram

Core XPath Expressions

//@class

— Select all attributes named

class
/article

— Select root element

article
//div

— Select all div child elements article — Select all child nodes of all article elements article/a — Select all a elements that are children of

article
article//div

— Select all div descendants of article Key distinction: / selects direct children only; // selects all descendants (broader scope). The @ prefix selects attributes (commonly @class).

Attribute-Based and Positional Selection

//div[@lang]

— Select all div elements that have a lang attribute //div[@lang='eng'] — Select all div elements with lang attribute equal to

eng
/article/div[1]

— Select the first div child of

article
/article/div[last()]

— Select the last div child of

article
/div/*

— Select all child nodes of div elements //* — Select all elements in the document //div/a | //div/p — Select all a and p elements that are children of div (union operator)

HTML source code showing class attribute
HTML source code showing class attribute

In the HTML source, class is an attribute name and the value after = (e.g., grid-5) is the attribute value; a node can have multiple attributes.

Practical Application

After mastering these XPath expressions, you can write queries to extract target data from web pages. The next article will cover using XPath in Scrapy spider projects.

XPath expression example
XPath expression example
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

html-parsingxmldata-extractionWeb ScrapingscrapyXPathnode-selection
Python Crawling & Data Mining
Written by

Python Crawling & Data Mining

Life's short, I code in Python. This channel shares Python web crawling, data mining, analysis, processing, visualization, automated testing, DevOps, big data, AI, cloud computing, machine learning tools, resources, news, technical articles, tutorial videos and learning materials. Join us!

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.