Skip to content

Utilities API

pathlib_next.utils

LRU(func, maxsize=128, on_evict=None)

Bases: Generic[K, V]

Thread-safe memoizing LRU cache over a function, callable like the function itself; invalidate(*args) evicts and recomputes an entry, discard(*args) evicts without recomputing.

on_evict(key, value), when given, is called for every value the cache drops -- LRU overflow, a maxsize shrink, invalidate()/discard(), and the losing result of two concurrent misses on one key -- so a cache of live resources (connections) can close them instead of leaking them. It runs outside the lock, and an exception it raises is suppressed: eviction is a side effect of an unrelated lookup, which must not fail because a stale resource would not close cleanly.

The computation itself runs outside the lock, so concurrent misses on one key may each call func; the first stored result wins and is returned to every caller, and the others are passed to on_evict.

Source code in src/pathlib_next/utils/__init__.py
58
59
60
61
62
63
64
65
66
67
68
def __init__(
    self,
    func: _ty.Callable[K, V],
    maxsize=128,
    on_evict: "_ty.Callable[[tuple, V], object] | None" = None,
):
    self.cache = collections.OrderedDict()
    self.func = func
    self._maxsize = maxsize
    self.lock = RLock()
    self.on_evict = on_evict

discard(*args)

Drop the entry for args (passing it to on_evict) without recomputing it. Returns whether an entry was present.

Source code in src/pathlib_next/utils/__init__.py
118
119
120
121
122
123
124
125
126
def discard(self, *args: K.args) -> bool:
    """Drop the entry for `args` (passing it to `on_evict`) without
    recomputing it. Returns whether an entry was present."""
    with self.lock:
        if args not in self.cache:
            return False
        value = self.cache.pop(args)
    self._dispose([(args, value)])
    return True

as_error_handler(ignore_error, *, default=False)

Normalize an ignore_error argument into a callable error policy.

Every ignore_error parameter in this library accepts either a bool or a callable, but the callables have deliberately different arities per call site (Path.rm() -> (error, path), Path.copy() -> (error), PathSyncer.sync() -> (error, source, target, event)). Unifying those arities would break existing callers, so this helper only normalizes the bool case and passes a supplied callable through untouched -- it is invoked with whatever arguments its own call site already uses.

None means "no policy supplied": it resolves to default (False for every current caller, i.e. raise on the first error), which preserves Path.copy(ignore_error=None)'s documented meaning.

Centralizing this keeps a fourth call site from drifting back into calling a bool (see PathSyncer.sync()'s symlink branch, which did exactly that and raised TypeError: 'bool' object is not callable).

Source code in src/pathlib_next/utils/__init__.py
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
def as_error_handler(
    ignore_error: _ty.Union[bool, _ty.Callable[..., bool], None],
    *,
    default: bool = False,
) -> _ty.Callable[..., bool]:
    """Normalize an `ignore_error` argument into a *callable* error policy.

    Every `ignore_error` parameter in this library accepts either a bool or
    a callable, but the callables have **deliberately different arities**
    per call site (`Path.rm()` -> `(error, path)`, `Path.copy()` ->
    `(error)`, `PathSyncer.sync()` -> `(error, source, target, event)`).
    Unifying those arities would break existing callers, so this helper only
    normalizes the *bool* case and passes a supplied callable through
    untouched -- it is invoked with whatever arguments its own call site
    already uses.

    `None` means "no policy supplied": it resolves to `default` (False for
    every current caller, i.e. raise on the first error), which preserves
    `Path.copy(ignore_error=None)`'s documented meaning.

    Centralizing this keeps a fourth call site from drifting back into
    calling a bool (see `PathSyncer.sync()`'s symlink branch, which did
    exactly that and raised `TypeError: 'bool' object is not callable`).
    """
    if callable(ignore_error):
        return ignore_error
    if ignore_error is None:
        ignore_error = default
    result = bool(ignore_error)
    return lambda *args, **kwargs: result

as_mode(mode)

Normalize a permission mode to an int, parsing str as octal.

chmod("0755") is the spelling everyone actually writes a mode in -- chmod(1), Ansible, Dockerfiles, every shell script -- and stdlib refuses it (TypeError: 'str' object cannot be interpreted as an integer). This library accepts it, which makes the base explicit and non-negotiable rather than leaving it to each call site.

Why base 8 is mandatory here, and never a plain int(): "0755" parsed as decimal is 755, which is 0o1363 -- a different and valid mode. Nothing would raise; the file would just end up with permissions nobody intended. That is exactly why stdlib declines strings, so the only safe way to accept them is to parse them one way, in one place.

Accepts an optional 0o/0O prefix. Anything outside [0-7] raises ValueError rather than being coerced -- a mode is not a number that happens to be written in octal, it is octal.

An int passes through untouched (including 0o755, which is an int by the time it gets here -- the literal is resolved by the parser, so chmod(0o755) and chmod("0755") agree).

Source code in src/pathlib_next/utils/__init__.py
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
def as_mode(mode: _ty.Union[int, str]) -> int:
    """Normalize a permission `mode` to an int, parsing `str` as **octal**.

    `chmod("0755")` is the spelling everyone actually writes a mode in --
    `chmod(1)`, Ansible, Dockerfiles, every shell script -- and stdlib
    refuses it (`TypeError: 'str' object cannot be interpreted as an
    integer`). This library accepts it, which makes the base explicit and
    non-negotiable rather than leaving it to each call site.

    **Why base 8 is mandatory here, and never a plain `int()`:** `"0755"`
    parsed as decimal is 755, which is `0o1363` -- a different *and valid*
    mode. Nothing would raise; the file would just end up with permissions
    nobody intended. That is exactly why stdlib declines strings, so the
    only safe way to accept them is to parse them one way, in one place.

    Accepts an optional `0o`/`0O` prefix. Anything outside `[0-7]` raises
    `ValueError` rather than being coerced -- a mode is not a number that
    happens to be written in octal, it is octal.

    An `int` passes through untouched (including `0o755`, which *is* an
    int by the time it gets here -- the literal is resolved by the parser,
    so `chmod(0o755)` and `chmod("0755")` agree).
    """
    if isinstance(mode, str):
        text = mode.strip()
        if text[:2].lower() == "0o":
            text = text[2:]
        if not text or any(character not in "01234567" for character in text):
            raise ValueError(f"invalid octal mode: {mode!r}")
        return int(text, 8)
    return _operator.index(mode)

as_owner(uid, gid)

Normalize a chown() uid/gid pair to canonical int | None.

None means "leave unchanged". -1 is accepted as an alias for it, since that is how os.chown spells the same thing and callers arriving from the stdlib reach for it out of habit.

The point of centralizing this is that every backend spells "unchanged" differently -- os.chown wants -1, SFTP setstat wants the field omitted from the attrs entirely, and other middlewares want None. If each scheme translated the caller's input itself, that is three chances for the semantics to disagree. Normalizing once on Path means a backend's _chown() receives an already-canonical pair and only has to convert to its own wire spelling.

A str is passed through as a name (shutil.chown accepts user and group names, and it is useful not to force a caller to resolve them) -- backends that cannot resolve names should say so rather than guess.

Source code in src/pathlib_next/utils/__init__.py
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
def as_owner(
    uid: _ty.Union[int, str, None], gid: _ty.Union[int, str, None]
) -> "_ty.Tuple[_ty.Optional[int], _ty.Optional[int]]":
    """Normalize a `chown()` uid/gid pair to canonical `int | None`.

    `None` means "leave unchanged". `-1` is accepted as an alias for it,
    since that is how `os.chown` spells the same thing and callers arriving
    from the stdlib reach for it out of habit.

    The point of centralizing this is that **every backend spells
    "unchanged" differently** -- `os.chown` wants `-1`, SFTP `setstat` wants
    the field omitted from the attrs entirely, and other middlewares want
    `None`. If each scheme translated the caller's input itself, that is
    three chances for the semantics to disagree. Normalizing once on `Path`
    means a backend's `_chown()` receives an already-canonical pair and only
    has to convert to its own wire spelling.

    A `str` is passed through as a *name* (`shutil.chown` accepts user and
    group names, and it is useful not to force a caller to resolve them) --
    backends that cannot resolve names should say so rather than guess.
    """

    def _one(value):
        if value is None or isinstance(value, str):
            return value
        value = _operator.index(value)
        return None if value == -1 else value

    return _one(uid), _one(gid)

is_safe_child_name(name, *, windows=False)

Whether name is a single path component that stays inside its parent.

Use this before joining a name that came from somewhere untrusted (a remote listing, an archive member, a sync source) onto a destination. Always rejected: a non-str, "", ".", "..", and any name containing / or NUL. With windows=True (see is_windows_flavoured) also rejected: \, : (a drive prefix such as D:x, or an NTFS alternate data stream), and names that are empty or ./.. once Windows strips trailing dots and spaces (".. ").

Source code in src/pathlib_next/utils/__init__.py
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
def is_safe_child_name(name: object, *, windows: bool = False) -> bool:
    """Whether `name` is a single path component that stays inside its parent.

    Use this before joining a name that came from somewhere untrusted (a
    remote listing, an archive member, a sync source) onto a destination.
    Always rejected: a non-`str`, `""`, `"."`, `".."`, and any name
    containing `/` or NUL. With `windows=True` (see `is_windows_flavoured`)
    also rejected: `\\`, `:` (a drive prefix such as `D:x`, or an NTFS
    alternate data stream), and names that are empty or `.`/`..` once
    Windows strips trailing dots and spaces (`".. "`).
    """
    if not isinstance(name, str) or name in ("", ".", "..") or "/" in name:
        return False
    if "\0" in name:
        return False
    if windows:
        if "\\" in name or ":" in name:
            return False
        if name.rstrip(" .") == "":
            return False
    return True

is_windows_flavoured(path)

Whether names joined onto path are interpreted with Windows rules.

True for a pathlib.PureWindowsPath (so LocalPath on Windows) and for a path whose filepath is one (FileUri on Windows). Everything else -- MemPath, remote schemes, POSIX local paths -- treats \ and : as ordinary filename characters.

Source code in src/pathlib_next/utils/__init__.py
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
def is_windows_flavoured(path: object) -> bool:
    """Whether names joined onto `path` are interpreted with Windows rules.

    True for a `pathlib.PureWindowsPath` (so `LocalPath` on Windows) and for
    a path whose `filepath` is one (`FileUri` on Windows). Everything else --
    `MemPath`, remote schemes, POSIX local paths -- treats `\\` and `:` as
    ordinary filename characters.
    """
    if isinstance(path, _PureWindowsPath):
        return True
    try:
        filepath = getattr(path, "filepath", None)
    except Exception:
        return False
    return isinstance(filepath, _PureWindowsPath)

parsedate(date)

Convert a date to UTC epoch seconds, independent of the host timezone.

  • str (RFC 1123, RFC 850 or asctime, as in Last-Modified): parsed with email.utils.parsedate_tz, and its zone offset applied. A string with no zone is read as UTC -- HTTP dates are GMT by definition.
  • int/float: already epoch seconds, returned unchanged.
  • struct_time/tuple: read as UTC (calendar.timegm), minus its tm_gmtoff (or a parsedate_tz-style 10th element) when one is set.
  • None or an unparseable string: 0.

Local time (time.mktime) is never involved: it shifted every HTTP/DAV st_mtime by the host's UTC offset and overflowed on Windows for dates near the epoch.

Source code in src/pathlib_next/utils/__init__.py
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
def parsedate(date: _ty.Union[str, _time.struct_time, tuple, int, float, None]):
    """Convert a date to UTC epoch seconds, independent of the host timezone.

    - `str` (RFC 1123, RFC 850 or asctime, as in `Last-Modified`): parsed
      with `email.utils.parsedate_tz`, and its zone offset applied. A string
      with no zone is read as UTC -- HTTP dates are GMT by definition.
    - `int`/`float`: already epoch seconds, returned unchanged.
    - `struct_time`/tuple: read as UTC (`calendar.timegm`), minus its
      `tm_gmtoff` (or a `parsedate_tz`-style 10th element) when one is set.
    - `None` or an unparseable string: `0`.

    Local time (`time.mktime`) is never involved: it shifted every HTTP/DAV
    `st_mtime` by the host's UTC offset and overflowed on Windows for dates
    near the epoch.
    """
    # Missing/unparseable dates yield epoch 0, not "now" -- a caller with no
    # Last-Modified header shouldn't have that read as "just modified" and
    # poison checksum/sync freshness comparisons.
    if date is None:
        return 0
    if isinstance(date, (int, float)):
        return date
    if isinstance(date, str):
        try:
            parsed = _parsedate_tz(date)
            if parsed is None:
                return 0
            return _calendar.timegm(parsed[:6]) - (parsed[9] or 0)
        except (TypeError, ValueError, IndexError, OverflowError):
            return 0
    offset = getattr(date, "tm_gmtoff", None)
    if offset is None and len(date) > 9:
        offset = date[9]
    return _calendar.timegm(tuple(date)[:6]) - (offset or 0)

pathlib_next.utils.glob

Filename globbing over any pathlib_next Path.

Modelled on CPython's pathlib glob selectors, reworked to run over anything that implements _scandir()/iterdir()/is_dir()/name like pathlib.Path (LocalPath, MemPath, every UriPath scheme).

NonRelativePatternError

Bases: NotImplementedError, ValueError

An absolute or anchored glob pattern. pathlib raises NotImplementedError("Non-relative patterns are unsupported"); this is also a ValueError, so either except clause catches it.

full_match(segments, pattern, case_sensitive)

Match segments against a glob pattern that may contain "" components (pathlib 3.13's PurePath.full_match semantics): a "" matches zero or more segments, except a trailing "" after other components, which needs at least one ("a/" does not match "a").

Runs as a set-of-states automaton, O(len(segments) * len(pattern)), so repeated "**" cannot backtrack exponentially.

Source code in src/pathlib_next/utils/glob.py
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
def full_match(segments: _ty.Sequence[str], pattern: str, case_sensitive: bool) -> bool:
    """Match `segments` against a glob pattern that may contain "**"
    components (pathlib 3.13's PurePath.full_match semantics): a "**" matches
    zero or more segments, except a trailing "**" after other components,
    which needs at least one ("a/**" does not match "a").

    Runs as a set-of-states automaton, O(len(segments) * len(pattern)), so
    repeated "**" cannot backtrack exponentially.
    """
    pats = pattern.split("/")
    pats = _collapse_recursive(p for i, p in enumerate(pats) if p or i == 0)
    segs = [s for i, s in enumerate(segments) if s or i == 0]
    end = len(pats)

    def closure(states: _ty.Set[int]) -> _ty.Set[int]:
        for i in sorted(states):
            # A non-trailing (or sole) "**" may match zero segments.
            if i < end and pats[i] == RECURSIVE and (i < end - 1 or end == 1):
                states.add(i + 1)
        return states

    states = closure({0})
    for seg in segs:
        advanced: _ty.Set[int] = set()
        for i in states:
            if i == end:
                continue
            pat = pats[i]
            if pat == RECURSIVE:
                advanced.update((i, i + 1))
            elif compile_pattern(pat, case_sensitive).match(seg):
                advanced.add(i + 1)
        if not advanced:
            return False
        states = closure(advanced)
    return end in states

glob(path, *, dironly=False, root_dir=None, recursive=False, include_hidden=False, case_sensitive=None, native=True, on_error=None, bound_loops=False)

Return an iterator which yields the paths matching a pathname pattern.

path is the pattern itself, as a path (e.g. UriPath("file:/x/**/*.py")). The pattern may contain simple shell-style wildcards a la fnmatch. Like the stdlib glob module, and unlike Path.glob(), hidden entries (names starting with a dot) are not matched by wildcards and not descended into by "**" unless include_hidden is true.

If recursive is true, the pattern '' will match any files and zero or more directories and subdirectories. recursive=None decides from the pattern itself, as Path.glob() does: a "" component enables it. native= follows the running interpreter's rules, or applies one rule on every version -- see parse_pattern().

Source code in src/pathlib_next/utils/glob.py
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
def glob(
    path: _Globable,
    *,
    dironly: bool = False,
    root_dir: _Globable | None = None,
    recursive: bool | None = False,
    include_hidden: bool = False,
    case_sensitive: bool | None = None,
    native: bool = True,
    on_error: "_ty.Callable[[OSError], None] | None" = None,
    bound_loops: bool = False,
) -> _ty.Iterable[_Globable]:
    """Return an iterator which yields the paths matching a pathname pattern.

    `path` is the pattern itself, as a path (e.g. `UriPath("file:/x/**/*.py")`).
    The pattern may contain simple shell-style wildcards a la fnmatch. Like
    the stdlib `glob` module, and unlike `Path.glob()`, hidden entries (names
    starting with a dot) are not matched by wildcards and not descended into
    by "**" unless `include_hidden` is true.

    If recursive is true, the pattern '**' will match any files and
    zero or more directories and subdirectories. `recursive=None` decides
    from the pattern itself, as `Path.glob()` does: a "**" component enables
    it. `native=` follows the running interpreter's rules, or applies one
    rule on every version -- see `parse_pattern()`.
    """
    segments = list(path.segments)
    if len(segments) > 1 and segments[-1] == "":
        segments.pop()
        if native and not _TRAILING_SLASH_SELECTS_DIRS:
            pass  # pathlib ignores a trailing separator before 3.11
        else:
            dironly = True
    if recursive is None:
        recursive = RECURSIVE in segments
    if native and not _PARTIAL_DOUBLESTAR_ALLOWED:
        for segment in segments:
            if RECURSIVE in segment and segment != RECURSIVE:
                raise ValueError(
                    "Invalid pattern: '**' can only be an entire path component"
                )
    first_wildcard = next(
        (i for i, seg in enumerate(segments) if WILDCARD_PATTERN.search(seg)),
        max(len(segments) - 1, 0),
    )
    base = path.with_segments(*segments[:first_wildcard])
    if root_dir is not None:
        base = root_dir / base
    parts = [part for part in segments[first_wildcard:] if part not in ("", ".")]
    if not parts:
        if base.is_dir() if dironly else base.exists():
            yield base
        return
    yield from select(
        base,
        parts,
        dironly=dironly,
        recursive=recursive,
        include_hidden=include_hidden,
        case_sensitive=case_sensitive,
        native=native,
        on_error=on_error,
        bound_loops=bound_loops,
    )

parse_pattern(pattern, *, native=True)

Split a relative glob pattern (a /-separated string or a Pathname) into its components, and report whether it ended with a separator (directories only).

Raises ValueError for an empty pattern and NonRelativePatternError for an absolute one, as pathlib.Path.glob() does. Empty and "." components are dropped.

native=True (the default) follows the running interpreter on the two rules pathlib changed mid-series: a trailing "/" is ignored before 3.11 (it selects directories only from 3.11), and a component that merely CONTAINS "" ("a") raises ValueError before 3.13. native=False applies one rule on every version instead -- a trailing "/" always selects directories only, "a**" is always a plain wildcard -- so a pattern gives the same answer on every interpreter and every backend.

Source code in src/pathlib_next/utils/glob.py
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def parse_pattern(
    pattern: "str | _ty.Any", *, native: bool = True
) -> _ty.Tuple[_ty.List[str], bool]:
    """Split a relative glob `pattern` (a `/`-separated string or a
    `Pathname`) into its components, and report whether it ended with a
    separator (directories only).

    Raises `ValueError` for an empty pattern and `NonRelativePatternError`
    for an absolute one, as `pathlib.Path.glob()` does. Empty and "."
    components are dropped.

    `native=True` (the default) follows the running interpreter on the two
    rules pathlib changed mid-series: a trailing "/" is ignored before 3.11
    (it selects directories only from 3.11), and a component that merely
    CONTAINS "**" ("a**") raises `ValueError` before 3.13. `native=False`
    applies one rule on every version instead -- a trailing "/" always
    selects directories only, "a**" is always a plain wildcard -- so a
    pattern gives the same answer on every interpreter and every backend.
    """
    if isinstance(pattern, str):
        rooted = pattern.startswith("/")
        segments = pattern.split("/")
    else:
        segments = list(pattern.segments)
        rooted = bool(getattr(pattern, "anchor", "")) or bool(
            getattr(pattern, "source", None)
        )
    if rooted:
        raise NonRelativePatternError("Non-relative patterns are unsupported")
    parts = [part for part in segments if part not in ("", ".")]
    if not parts:
        raise ValueError(f"Unacceptable pattern: {str(pattern)!r}")
    if native and not _PARTIAL_DOUBLESTAR_ALLOWED:
        for part in parts:
            if RECURSIVE in part and part != RECURSIVE:
                # pathlib's own message, so an `except ValueError` that reads
                # it sees the same text it does on this interpreter.
                raise ValueError(
                    "Invalid pattern: '**' can only be an entire path component"
                )
    trailing_sep = bool(segments) and segments[-1] == ""
    if native and not _TRAILING_SLASH_SELECTS_DIRS:
        trailing_sep = False
    return parts, trailing_sep

select(base, parts, *, dironly=False, recursive=True, include_hidden=True, case_sensitive=None, native=True, on_error=None, bound_loops=False)

Yield the paths under base matching the pattern components parts (see parse_pattern()); the engine behind Path.glob().

Directory listings go through _scandir(); an OSError from listing a missing or non-directory path selects nothing. "" (when recursive) decides recursion from the listing's non-following stat, so it never descends into a directory symlink and always terminates. A trailing "" also selects files on Python 3.13+ and directories only before, which is the third rule native=False pins: it then selects files on every version (3.13's rule, the one current pathlib applies).

Source code in src/pathlib_next/utils/glob.py
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
def select(
    base: _Globable,
    parts: _ty.Sequence[str],
    *,
    dironly: bool = False,
    recursive: bool = True,
    include_hidden: bool = True,
    case_sensitive: bool | None = None,
    native: bool = True,
    on_error: "_ty.Callable[[OSError], None] | None" = None,
    bound_loops: bool = False,
) -> _ty.Iterator[_Globable]:
    """Yield the paths under `base` matching the pattern components `parts`
    (see `parse_pattern()`); the engine behind `Path.glob()`.

    Directory listings go through `_scandir()`; an `OSError` from listing a
    missing or non-directory path selects nothing. "**" (when `recursive`)
    decides recursion from the listing's non-following stat, so it never
    descends into a directory symlink and always terminates. A trailing "**"
    also selects files on Python 3.13+ and directories only before, which is
    the third rule `native=False` pins: it then selects files on every
    version (3.13's rule, the one current pathlib applies).
    """
    default_case = getattr(base, "_is_case_sensitive", True)
    if case_sensitive is None:
        case_sensitive = default_case
    if recursive:
        parts = _collapse_recursive(parts)
    steps: _ty.List[_ty.Tuple[str, _ty.Any]] = []
    for part in parts:
        if recursive and part == RECURSIVE:
            steps.append((part, None))
        elif WILDCARD_PATTERN.search(part) or case_sensitive != default_case:
            steps.append((part, compile_pattern(part, case_sensitive)))
        else:
            steps.append((part, False))
    opts = _Options(
        dironly,
        include_hidden,
        _DOUBLESTAR_SELECTS_FILES if native else True,
        on_error,
        bound_loops,
    )
    selected = _select(base, steps, 0, opts, None)
    if sum(1 for _, kind in steps if kind is None) < 2:
        yield from selected
        return
    # Two separate "**" can reach one path along several splits.
    seen = set()
    for path in selected:
        if path not in seen:
            seen.add(path)
            yield path

pathlib_next.utils.sync

PathSyncer(checksum=None, /, remove_missing=False, follow_symlinks=True, symlink_mode='preserve', hook=None, ignore_error=False, quick_check=True)

Bases: object

One-way checksum-driven tree sync: copies/creates in target whatever differs from source (by checksum), optionally removing files in target that are missing from source. Works across any two Path implementations (e.g. MemPath -> LocalPath, or between two UriPath schemes) -- see sync().

The default checksum policy prefers each side's backend-native digest (protocols.checksum.NativeChecksum.checksum(), e.g. SftpPath's check-file-handle support) over streaming the file through open("rb"), but only when BOTH sides can produce a digest under the same algorithm -- native or streamed. If either side can't (missing the protocol, or it raises NotImplementedError for the requested algorithm), both sides fall back to streaming rather than comparing a native digest to a streamed one. A custom checksum callable disables this native-preferring behavior entirely (it is called exactly as before, once per side, compared with ==).

quick_check=True (the default) adds a cheap metadata-only pre-check -- the classic rsync "quick check" heuristic -- for any pair where at least one side is non-local (Uri.is_local(); a side without an is_local() method at all, e.g. plain LocalPath/MemPath, is treated as local): if st_size AND st_mtime already match (from the listing/stat metadata PathAndStat already carries -- no extra round trip), the pair is treated as in sync WITHOUT calling checksum at all, native or streamed. A mismatch on either falls through to a real checksum comparison rather than being treated as "changed" -- mtime can be unreliable across backends/clock skew, so a false "needs copy" from a mismatch is merely wasteful, while a false "in sync" would be a correctness regression. An st_mtime of 0 on either side means "unknown" and never matches. Local-to-local pairs always skip this pre-check (unchanged pre-existing behavior -- local reads are already cheap, and this project's copy(preserve_metadata=True) doesn't guarantee mtime propagation on every path, see docs/divergences.md). Set quick_check=False to disable the pre-check entirely and always checksum, matching pre-quick_check behavior for non-local pairs too.

follow_symlinks (default True) controls whether a symlink source is resolved during traversal (content synced as if it weren't a link) or reported as a symlink (is_symlink() true). When it's False and a symlink source is reached, symlink_mode decides what happens: "preserve" (default) creates a matching symlink on target with the same raw, unresolved target string readlink() returned (dangling links and relative targets included -- never resolved/validated); "reject" raises NotImplementedError instead (the only behavior before this kwarg existed). If target can't create symlinks at all (most backends -- only LocalPath and SftpPath currently implement symlink_to()), "preserve" mode raises NotImplementedError too, through the same ignore_error/hook() machinery as every other branch -- decided before an existing target entry is touched. Replacing an existing entry creates the new link under a temporary sibling name first, so a runtime refusal also leaves the entry in place. A target link that already has the same raw target string (and, on a Windows target, the same file/directory link kind) is left alone. A directory link is created as one (target_is_directory), which Windows requires.

remove_missing=False never deletes a non-empty target directory: when the source entry became a file or a symlink, replacing that directory raises IsADirectoryError through ignore_error (event SyncEvent.TypeMismatch) and the directory is left untouched. With remove_missing=True, or for an empty directory, it is replaced.

A changed file is written to a temporary sibling and renamed over the existing target where the target backend implements rename(), so a failed or interrupted transfer keeps the previous version. A backend without rename() is overwritten in place, with the source opened before the target is truncated. A source entry that is not a regular file, directory or symlink (FIFO, socket, device) is skipped with a SyncEvent.Skipped event and its target left untouched.

Errors: ignore_error is consulted once per error, with the paths of the entry that failed and the event that failed (SyncEvent.Compare for a failing quick check or checksum). An error it declines propagates without being offered again by enclosing directories. A dry run takes the same decisions as a real run (a directory that would be created is treated as empty) without changing anything. Every error the policy tolerates is logged at WARNING (with the traceback) on the pathlib_next.sync logger and reported to hook as SyncEvent.Error.

hook receives PathAndStat entries for every event, and the dry_run value of the call for every event, including the structural ones (SyncStart, CheckTargetChild(ren), SyncChild(ren)).

Source code in src/pathlib_next/utils/sync.py
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
def __init__(
    self,
    checksum: _ty.Callable[[PathAndStat], _ty.Any] | None = None,
    /,
    remove_missing: bool = False,
    follow_symlinks: bool = True,
    symlink_mode: '_ty.Literal["preserve", "reject"]' = "preserve",
    hook: _ty.Callable[[PathAndStat, PathAndStat, SyncEvent, bool], None] = None,
    ignore_error: _OnPathSyncerError | bool = False,
    quick_check: bool = True,
) -> None:
    # `None` (the default) resolves to `_default_checksum` -- a sentinel
    # `sync()` recognizes (via `is`) to route through
    # `_default_checksums_match()` instead of two independent calls, so
    # the native-vs-streaming decision can be coordinated across BOTH
    # sides at once. A caller-supplied callable is stored and used
    # as-is (`checksum(target) == checksum(source)`, unchanged from
    # before this feature).
    if checksum is None:
        checksum = _default_checksum
    self.checksum = checksum
    self.remove_missing = remove_missing
    self._hook = hook
    self.follow_symlinks = follow_symlinks
    if symlink_mode not in ("preserve", "reject"):
        raise ValueError(
            f"symlink_mode must be 'preserve' or 'reject', got {symlink_mode!r}"
        )
    self.symlink_mode = symlink_mode
    self.ignore_error = _ty.cast(
        _OnPathSyncerError, _utils.as_error_handler(ignore_error)
    )
    self.quick_check = quick_check

sync(source, target, /, dry_run=False, ignore_error=None)

Sync source onto target.

ignore_error overrides the instance-level policy for this call only. It accepts a bool or a callable with the same (error, source, target, event) arity as the constructor's; None (the default) means "use the policy given to __init__".

The default used to be the bool False, which both shadowed a constructor-supplied policy and was called directly by the symlink branch (TypeError: 'bool' object is not callable). Passing a callable explicitly behaves exactly as before.

The root call is checked before anything is touched: a source that does not exist raises FileNotFoundError (only a child that vanishes mid-sync takes the remove_missing path), and a source and target of the same implementation where one contains the other raise ValueError. Both go through ignore_error; a tolerated error ends the call without changes. Child names that would not stay a single component inside target (.., a separator or drive on a Windows target) raise ValueError the same way, and symlinks found inside target are replaced, never written, listed or deleted through.

Source code in src/pathlib_next/utils/sync.py
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
def sync(
    self,
    source: Path | PathAndStat,
    target: Path | PathAndStat,
    /,
    dry_run: bool = False,
    ignore_error: _OnPathSyncerError | bool | None = None,
):
    """Sync `source` onto `target`.

    `ignore_error` overrides the instance-level policy for this call
    only. It accepts a bool or a callable with the same
    `(error, source, target, event)` arity as the constructor's; `None`
    (the default) means "use the policy given to `__init__`".

    The default used to be the bool `False`, which both shadowed a
    constructor-supplied policy and was *called* directly by the symlink
    branch (`TypeError: 'bool' object is not callable`). Passing a
    callable explicitly behaves exactly as before.

    The root call is checked before anything is touched: a `source`
    that does not exist raises `FileNotFoundError` (only a child that
    vanishes mid-sync takes the `remove_missing` path), and a `source`
    and `target` of the same implementation where one contains the
    other raise `ValueError`. Both go through `ignore_error`; a
    tolerated error ends the call without changes. Child names that
    would not stay a single component inside `target` (`..`, a
    separator or drive on a Windows target) raise `ValueError` the
    same way, and symlinks found inside `target` are replaced, never
    written, listed or deleted through.
    """
    _ignore_error = (
        self.ignore_error
        if ignore_error is None
        else _utils.as_error_handler(ignore_error)
    )
    return self._sync(source, target, dry_run, _ignore_error, True)

SyncEvent

Bases: Enum

Events PathSyncer.hook() fires during a sync, for progress/logging callbacks.

pathlib_next.utils.stat

FileStat(st_mode=None, st_size=0, st_mtime=0, is_dir=False)

Bases: FileStatLike

Concrete, slotted FileStatLike for backends without a real os.stat_result (e.g. MemPath, HttpPath). from_path() builds one from any object with a stat() method, or passes a FileStat through unchanged.

Source code in src/pathlib_next/utils/stat.py
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
def __init__(
    self,
    st_mode: int = None,
    st_size: int = 0,
    st_mtime: int = 0,
    is_dir: bool = False,
):
    self.mode_known = bool(st_mode)
    self.st_mode = st_mode or (
        _stat.S_IFDIR | 0o555 if is_dir else _stat.S_IFREG | 0o444
    )
    self.st_nlink = 1
    self.st_uid = 0
    self.st_gid = 0
    self.st_size = st_size
    self.st_atime = 0
    self.st_mtime = st_mtime
    self.st_ctime = 0

from_stat(stat) classmethod

Copy any stat-like object's (os.stat_result, paramiko's SFTPAttributes, ...) recognized fields into a fresh FileStat, so downstream code (e.g. .is_dir()) can rely on a uniform type. Passes an already-FileStat through unchanged. The copied mode counts as backend-reported (mode_known) unless the source says otherwise or carries no mode at all (missing, None or 0).

A missing or None field becomes 0: paramiko's SFTPAttributes leaves every field the server did not send as None, which made is_dir() and repr() raise TypeError.

Source code in src/pathlib_next/utils/stat.py
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
@classmethod
def from_stat(cls, stat: _ty.Any) -> "FileStat":
    """Copy any stat-like object's (`os.stat_result`, paramiko's
    `SFTPAttributes`, ...) recognized fields into a fresh `FileStat`,
    so downstream code (e.g. `.is_dir()`) can rely on a uniform type.
    Passes an already-`FileStat` through unchanged. The copied mode
    counts as backend-reported (`mode_known`) unless the source says
    otherwise or carries no mode at all (missing, `None` or `0`).

    A missing or `None` field becomes `0`: paramiko's `SFTPAttributes`
    leaves every field the server did not send as `None`, which made
    `is_dir()` and `repr()` raise TypeError."""
    if isinstance(stat, FileStat):
        return stat
    result = FileStat.__new__(FileStat)
    for prop in FileStat._FIELDS:
        value = getattr(stat, prop, None)
        setattr(result, prop, 0 if value is None else value)
    result.mode_known = bool(getattr(stat, "mode_known", True)) and bool(
        result.st_mode
    )
    return result

is_block_device()

Whether this path is a block device.

Source code in src/pathlib_next/utils/stat.py
142
143
144
145
146
def is_block_device(self):
    """
    Whether this path is a block device.
    """
    return _stat.S_ISBLK(self.st_mode)

is_char_device()

Whether this path is a character device.

Source code in src/pathlib_next/utils/stat.py
148
149
150
151
152
def is_char_device(self):
    """
    Whether this path is a character device.
    """
    return _stat.S_ISCHR(self.st_mode)

is_dir()

Whether this path is a directory.

Source code in src/pathlib_next/utils/stat.py
123
124
125
126
127
def is_dir(self):
    """
    Whether this path is a directory.
    """
    return _stat.S_ISDIR(self.st_mode)

is_fifo()

Whether this path is a FIFO.

Source code in src/pathlib_next/utils/stat.py
154
155
156
157
158
def is_fifo(self):
    """
    Whether this path is a FIFO.
    """
    return _stat.S_ISFIFO(self.st_mode)

is_file()

Whether this path is a regular file (also True for symlinks pointing to regular files).

Source code in src/pathlib_next/utils/stat.py
129
130
131
132
133
134
def is_file(self):
    """
    Whether this path is a regular file (also True for symlinks pointing
    to regular files).
    """
    return _stat.S_ISREG(self.st_mode)

is_socket()

Whether this path is a socket.

Source code in src/pathlib_next/utils/stat.py
160
161
162
163
164
def is_socket(self):
    """
    Whether this path is a socket.
    """
    return _stat.S_ISSOCK(self.st_mode)

Whether this path is a symbolic link.

Source code in src/pathlib_next/utils/stat.py
136
137
138
139
140
def is_symlink(self):
    """
    Whether this path is a symbolic link.
    """
    return _stat.S_ISLNK(self.st_mode)

pathlib_next.utils.checksum

md5(path, chunk_size=65536)

Calculate MD5 checksum of the file at path.

Source code in src/pathlib_next/utils/checksum.py
27
28
29
def md5(path: Path, chunk_size: int = 65536) -> str:
    """Calculate MD5 checksum of the file at `path`."""
    return stream(path, "md5", chunk_size)

native(path, algorithm='md5')

Try path's backend-native NativeChecksum.checksum() (see protocols/checksum.py); return None if path doesn't implement the protocol at all, or if it does but can't produce algorithm (raises NotImplementedError -- the protocol's documented contract for "unsupported", never a wrong-algorithm value). Never raises for either of those two cases; any other exception the backend raises propagates.

Source code in src/pathlib_next/utils/checksum.py
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
def native(path: "Path", algorithm: str = "md5") -> "str | None":
    """Try `path`'s backend-native `NativeChecksum.checksum()` (see
    `protocols/checksum.py`); return `None` if `path` doesn't implement the
    protocol at all, or if it does but can't produce `algorithm` (raises
    `NotImplementedError` -- the protocol's documented contract for
    "unsupported", never a wrong-algorithm value). Never raises for either
    of those two cases; any other exception the backend raises propagates.
    """
    checksum = getattr(path, "checksum", None)
    if checksum is None:
        return None
    try:
        return checksum(algorithm)
    except NotImplementedError:
        return None

sha256(path, chunk_size=65536)

Calculate SHA-256 checksum of the file at path.

Source code in src/pathlib_next/utils/checksum.py
32
33
34
def sha256(path: Path, chunk_size: int = 65536) -> str:
    """Calculate SHA-256 checksum of the file at `path`."""
    return stream(path, "sha256", chunk_size)

stream(path, algorithm='md5', chunk_size=65536)

Streaming checksum for an arbitrary hashlib algorithm name (the generic form of md5/sha256 above, used where the algorithm is a runtime parameter rather than fixed at the call site -- e.g. PathSyncer's streaming fallback, which must match whatever algorithm a native checksum attempt was made under).

The digest compares content; it is not a security use, so it is created with usedforsecurity=False -- on a FIPS-mode host hashlib.md5() otherwise raises ValueError.

Source code in src/pathlib_next/utils/checksum.py
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
def stream(path: "Path", algorithm: str = "md5", chunk_size: int = 65536) -> str:
    """Streaming checksum for an arbitrary `hashlib` `algorithm` name (the
    generic form of `md5`/`sha256` above, used where the algorithm is a
    runtime parameter rather than fixed at the call site -- e.g.
    `PathSyncer`'s streaming fallback, which must match whatever algorithm
    a native checksum attempt was made under).

    The digest compares content; it is not a security use, so it is created
    with `usedforsecurity=False` -- on a FIPS-mode host `hashlib.md5()`
    otherwise raises ValueError.
    """
    h = _hashlib.new(algorithm, usedforsecurity=False)
    with path.open("rb") as f:
        while True:
            chunk = f.read(chunk_size)
            if not chunk:
                break
            h.update(chunk)
    return h.hexdigest()

pathlib_next.utils.archive

make_archive(src, format, target)

Create an archive file from src at target.

Supports format='zip' and format='tar'. Operations run stream-first to support any Path implementation, for src as well as target. The archive is built in a temporary buffer and written to target only once complete, so a failure leaves an existing target untouched. Zip members always carry zip64 headers, so a member over 2 GiB is stored instead of failing after the fact.

Source code in src/pathlib_next/utils/archive.py
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def make_archive(src: Path, format: str, target: Path) -> None:
    """Create an archive file from `src` at `target`.

    Supports format='zip' and format='tar'.
    Operations run stream-first to support any Path implementation, for
    `src` as well as `target`. The archive is built in a temporary buffer
    and written to `target` only once complete, so a failure leaves an
    existing `target` untouched. Zip members always carry zip64 headers,
    so a member over 2 GiB is stored instead of failing after the fact.
    """
    if format not in ("zip", "tar"):
        raise ValueError(f"Unsupported format: {format}")

    if src.is_file():
        members = [(src, src.name)]
    elif src.is_dir():
        members = _archive_members(src)
    else:
        raise FileNotFoundError(f"No such file or directory: {str(src)!r}")

    with _Spool() as buffer:
        if format == "zip":
            with zipfile.ZipFile(buffer, "w", zipfile.ZIP_DEFLATED) as archive:
                for file_path, arcname in members:
                    # force_zip64: the size is unknown when the entry opens,
                    # and without it zipfile raises once 2 GiB were written.
                    with archive.open(arcname, "w", force_zip64=True) as dest_f:
                        with file_path.open("rb") as src_f:
                            _copy_chunks(src_f, dest_f)
        else:
            with tarfile.open(fileobj=buffer, mode="w") as archive:
                for file_path, arcname in members:
                    stat = file_path.stat()
                    info = tarfile.TarInfo(name=arcname)
                    info.size = stat.st_size
                    info.mode = stat.st_mode
                    with file_path.open("rb") as src_f:
                        archive.addfile(info, src_f)
        buffer.seek(0)
        with target.open("wb") as out_f:
            shutil.copyfileobj(buffer, out_f)

unpack_archive(archive, dest)

Extract archive file into dest directory.

Supports format detection from filename. Operations run stream-first to support any Path implementation. Members whose name would land outside dest (a .. part, or on a Windows-flavoured dest a drive such as D:x or C:..) are skipped. A non-seekable archive stream (e.g. HttpPath) is buffered first. A tar hard link, or a symlink to a regular file inside the archive, is extracted as a regular file with the target member's content (no link is created, so dest needs no symlink support); a link that does not resolve to such a member is skipped with a UserWarning.

Source code in src/pathlib_next/utils/archive.py
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
def unpack_archive(archive: Path, dest: Path) -> None:
    """Extract `archive` file into `dest` directory.

    Supports format detection from filename.
    Operations run stream-first to support any Path implementation.
    Members whose name would land outside `dest` (a `..` part, or on a
    Windows-flavoured `dest` a drive such as `D:x` or `C:..`) are skipped.
    A non-seekable archive stream (e.g. `HttpPath`) is buffered first.
    A tar hard link, or a symlink to a regular file inside the archive, is
    extracted as a regular file with the target member's content (no link
    is created, so `dest` needs no symlink support); a link that does not
    resolve to such a member is skipped with a `UserWarning`.
    """
    if not dest.exists():
        dest.mkdir(parents=True, exist_ok=True)

    def _peek() -> bytes:
        try:
            with archive.open("rb") as f:
                return f.read(4)
        except Exception:
            return b"PK"  # can't sniff -- preserve the historical "assume zip" default

    is_zip = _detect_format(archive.name, _peek) == "zip"
    windows = is_windows_flavoured(dest)

    with archive.open("rb") as raw_in, _Spool() as spool:
        in_f = raw_in
        seekable = getattr(raw_in, "seekable", None)
        if seekable is None or not seekable():
            # zipfile seeks to the central directory, and tarfile's "r:*"
            # probe seeks back: neither works over a network stream.
            shutil.copyfileobj(raw_in, spool)
            spool.seek(0)
            in_f = spool
        if is_zip:
            with zipfile.ZipFile(in_f) as zip_ref:
                for member in zip_ref.infolist():
                    filename = member.filename
                    parts = _safe_member_parts(filename, windows=windows)
                    if parts is None:
                        continue

                    target_path = dest
                    for part in parts:
                        target_path = target_path / part

                    if member.is_dir() or filename.endswith("/"):
                        target_path.mkdir(parents=True, exist_ok=True)
                    else:
                        target_path.parent.mkdir(parents=True, exist_ok=True)
                        with target_path.open("wb") as out_f:
                            with zip_ref.open(member) as member_f:
                                _copy_chunks(member_f, out_f)
        else:
            with tarfile.open(fileobj=in_f, mode="r") as tar_ref:
                for member in tar_ref.getmembers():
                    filename = member.name
                    parts = _safe_member_parts(filename, windows=windows)
                    if parts is None:
                        continue

                    target_path = dest
                    for part in parts:
                        target_path = target_path / part

                    if member.isdir():
                        target_path.mkdir(parents=True, exist_ok=True)
                    elif member.isfile():
                        target_path.parent.mkdir(parents=True, exist_ok=True)
                        with target_path.open("wb") as out_f:
                            member_f = tar_ref.extractfile(member)
                            if member_f is not None:
                                _copy_chunks(member_f, out_f)
                    elif member.islnk() or member.issym():
                        # extractfile() resolves a link to the member it
                        # names inside the archive -- never the filesystem.
                        try:
                            member_f = tar_ref.extractfile(member)
                        except KeyError:
                            member_f = None
                        if member_f is None:
                            warnings.warn(
                                f"unpack_archive: skipped link {filename!r} -> "
                                f"{member.linkname!r}: not a regular file in "
                                "the archive",
                                UserWarning,
                                stacklevel=2,
                            )
                            continue
                        target_path.parent.mkdir(parents=True, exist_ok=True)
                        with member_f, target_path.open("wb") as out_f:
                            _copy_chunks(member_f, out_f)