ScreenshotNeo

BlogHow-to

How to Select All Elements Between Two Elements in XPath

Learn XPath patterns for selecting nodes between start and end markers, excluding or including boundaries, repeated sections, namespaces, and XPath versions.

By the ScreenshotNeo team1 October 20266 min read

Use this XPath for elements with the same parent:

//item[preceding-sibling::start and following-sibling::end]

This returns item elements that have a start sibling before them and an end sibling after them. The two boundary elements are excluded.

1. Select siblings strictly between two markers

Given this XML:

<section>
  <start id='one'/>
  <item>A</item>
  <item>B</item>
  <end id='one'/>
</section>

The expression below selects both item elements:

//item[preceding-sibling::start and following-sibling::end]

preceding-sibling::start checks for a matching sibling before the current node. following-sibling::end checks for a matching sibling after it. Because both tests must be true, the markers themselves do not match.

Select any element between the markers

//*[preceding-sibling::start and following-sibling::end]

Use * when the nodes between the markers can have different element names.

Constrain the element and marker attributes

//div[@class='entry'][preceding-sibling::h2[@id='start'] and following-sibling::h2[@id='end']]

This is safer when the document contains unrelated start or end elements.

2. Include one or both boundary elements

The between predicate excludes both boundaries. Add a union when the result must include them.

//start | //item[preceding-sibling::start and following-sibling::end] | //end

Include only the start marker:

//start | //item[preceding-sibling::start and following-sibling::end]

Include only the end marker:

//item[preceding-sibling::start and following-sibling::end] | //end

Wrap a union in parentheses before applying a position:

(//start | //item[preceding-sibling::start and following-sibling::end] | //end)[1]

3. Select nodes between markers in different branches

preceding-sibling and following-sibling work only when the boundaries share a parent with the target nodes. If the markers are in different branches, use document-order axes.

In XPath 1.0, this intersection pattern selects nodes after the first marker and before the second:

(//incision[2]/preceding::*)[
  count(. | (//incision[1]/following::*))
  = count((//incision[1]/following::*))
]

Replace incision and the occurrence numbers with your marker name and intended instances. The expression keeps nodes that occur in both sets: nodes before the second marker and nodes after the first.

The following axis contains nodes after the context node in document order, excluding its descendants. The preceding axis contains nodes before the context node, excluding its ancestors. This distinction matters when a marker contains nested content.

XPath 2.0 and later

XPath 2.0+ host APIs can bind the two marker nodes and compare node order or filter sequences. The exact syntax depends on the processor, so check whether your library exposes XPath 1.0, XPath 2.0, XPath 3.1, or XQuery. Browser DOM APIs generally expose XPath 1.0.

4. Handle repeated sections and nearest boundaries

A basic predicate can span farther than intended when a document has multiple sections. For example, an item in the second section may still have some earlier start sibling in the same parent.

Require the nearest matching boundary and identify it explicitly:

//item[
  preceding-sibling::start[1][@id='start-1']
  and following-sibling::end[1][@id='end-1']
]

For document-wide markers, select their occurrences before applying the relationship:

(//start)[1]
(//end)[1]

Use explicit IDs, section containers, or occurrence indexes whenever markers repeat. Test nested sections separately: a flat between test does not model overlapping or nested ranges.

5. Select text, comments, or other node kinds

The * node test selects element nodes. Use node() when the result must include text nodes, comments, or processing instructions:

//node()[preceding-sibling::start and following-sibling::end]

Attributes and namespace nodes are not ordinary child elements. Select attributes with an attribute axis, such as //item/@data-id. Select namespace nodes only when your XPath processor supports that axis.

6. Namespaces and evaluation context

Namespace-qualified XML requires a namespace binding in the host language. A visible prefix in the XML is not automatically available to the XPath evaluator.

//svg:rect[preceding-sibling::svg:start and following-sibling::svg:end]

Bind svg to the document’s namespace URI in your XML library. If your API uses expanded names, use that API rather than matching an unbound literal prefix.

Also verify the context node. A relative path such as item[preceding-sibling::start] searches from the current context, while a leading // searches descendants from the document context.

7. Complete examples in common environments

Browser JavaScript (XPath 1.0)

const xml = `
A B
`; const doc = new DOMParser().parseFromString(xml, 'application/xml'); const xpath = "//item[preceding-sibling::start and following-sibling::end]"; const result = doc.evaluate(xpath, doc, null, XPathResult.ORDERED_NODE_SNAPSHOT_TYPE, null); const items = []; for (let i = 0; i < result.snapshotLength; i++) { items.push(result.snapshotItem(i).textContent.trim()); } console.log(items); // ['A', 'B']

Python with lxml

from lxml import etree

xml = '''<section>
  <start id="one"/>
  <item>A</item>
  <item>B</item>
  <end id="one"/>
</section>'''

root = etree.fromstring(xml.encode())
items = root.xpath(".//item[preceding-sibling::start and following-sibling::end]")
print([item.text for item in items])

Java with XPath

XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodes = (NodeList) xpath.evaluate(
    "//item[preceding-sibling::start and following-sibling::end]",
    document,
    XPathConstants.NODESET
);

8. A practical checklist

  1. Confirm whether both markers share the target node’s parent.
  2. Use preceding-sibling and following-sibling for same-parent ranges.
  3. Use following and preceding for different branches.
  4. Decide whether the boundaries should be excluded or added with a union.
  5. Qualify repeated markers by ID, occurrence, container, or nearest-marker logic.
  6. Choose * for elements or node() for all node kinds.
  7. Bind namespaces in the host API.
  8. Check the XPath version supported by the evaluator.
  9. Test empty ranges, missing markers, repeated markers, and nested sections.

9. Troubleshooting

No nodes are returned

Check the context node, spelling, namespace binding, and whether the markers are actually siblings. In browser APIs, an unbound namespace prefix commonly produces an empty result.

The result includes nodes from another section

The marker test is too broad. Add an ID, section container, occurrence index, or nearest-boundary predicate such as preceding-sibling::start[1].

The markers are missing from the result

That is expected for the strict-between expression. Add the desired boundary with a union.

Nested content is unexpectedly included or excluded

Sibling axes inspect only children of the same parent. For branch-wide ranges, use document-order axes and decide whether descendants of a marker should count.

The XPath works in one tool but not another

Compare XPath versions and result APIs. Browser document.evaluate is XPath 1.0; XPath 2.0+ operators and sequence features may not be available.

Only elements are returned

Replace * with node() when text, comments, or processing instructions are required.

10. Performance and reliability

Keep the search scope as narrow as possible. Starting from a section container, using specific marker names, and adding IDs reduces work compared with several document-wide // searches.

For large XML documents, parse once, compile reusable XPath expressions when your library supports it, and avoid repeatedly evaluating broad document-order axes. If the document can change between evaluations, keep the marker and result selection in the same snapshot or transaction.

Validate the input before evaluating XPath. Malformed XML, unexpected namespaces, and missing markers should produce a clear application error rather than an empty result that looks valid.

11. Or skip the browser setup

If your goal is to inspect a live page before selecting content, ScreenshotNeo can capture the page in one request. See the API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. ScreenshotNeo also provides an MCP server so Claude, Cursor, and other MCP clients can take screenshots, inspect page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Start with 1,000 free screenshots a month.

12. FAQ

What is the shortest XPath for same-parent boundaries?

//item[preceding-sibling::start and following-sibling::end].

How do I select everything between the second start and second end?

Qualify the intended occurrences, for example //item[preceding-sibling::start[@id='s2'] and following-sibling::end[@id='e2']], or isolate the relevant section first.

Can XPath select a range by position?

Position predicates select nodes in an axis result, but they do not reliably express two changing boundaries. Marker predicates are clearer when the document structure is semantic.

Does this work for HTML?

Yes, when the HTML parser exposes a DOM and XPath evaluator. Account for browser namespaces and parser error recovery, especially with SVG or MathML.

How do I include whitespace text nodes?

Use node() and handle whitespace in application code, because formatted XML commonly contains indentation text nodes.