Skip to content

ccdf

champalimaud.ccdf

The complementary cumulative distribution function (CCDF).

The CCDF of a sample is the fraction of the sample at or above a mark.

ccdf(sample)

Compute the CCDF of a sample at each of its distinct values.

Parameters:

Name Type Description Default
sample array_like

The values; at least one.

required

Returns:

Type Description
DataFrame

Columns value and fraction: for each distinct value, the fraction of the sample at or above it.

Examples:

>>> ccdf([1, 2, 2, 4]).rows()
[(1, 1.0), (2, 0.75), (4, 0.25)]
Source code in champalimaud/ccdf.py
def ccdf(sample: ArrayLike) -> pl.DataFrame:
    """Compute the CCDF of a sample at each of its distinct values.

    Parameters
    ----------
    sample : array_like
        The values; at least one.

    Returns
    -------
    polars.DataFrame
        Columns ``value`` and ``fraction``: for each distinct value,
        the fraction of the sample at or above it.

    Examples
    --------
    >>> ccdf([1, 2, 2, 4]).rows()
    [(1, 1.0), (2, 0.75), (4, 0.25)]
    """
    distinct = np.unique(np.asarray(sample))
    return pl.DataFrame(
        {"value": distinct, "fraction": ccdf_at(sample, distinct)}
    )

ccdf_at(sample, marks)

Evaluate the CCDF of a sample at some marks.

Parameters:

Name Type Description Default
sample array_like

The values; at least one.

required
marks array_like

Where to evaluate.

required

Returns:

Type Description
ndarray

For each mark, the fraction of sample at or above it.

Raises:

Type Description
ValueError

For an empty sample.

Examples:

>>> ccdf_at([1, 2, 2, 4], [1, 2, 3, 5]).tolist()
[1.0, 0.75, 0.25, 0.0]
Source code in champalimaud/ccdf.py
def ccdf_at(sample: ArrayLike, marks: ArrayLike) -> np.ndarray:
    """Evaluate the CCDF of a sample at some marks.

    Parameters
    ----------
    sample : array_like
        The values; at least one.
    marks : array_like
        Where to evaluate.

    Returns
    -------
    numpy.ndarray
        For each mark, the fraction of `sample` at or above it.

    Raises
    ------
    ValueError
        For an empty sample.

    Examples
    --------
    >>> ccdf_at([1, 2, 2, 4], [1, 2, 3, 5]).tolist()
    [1.0, 0.75, 0.25, 0.0]
    """
    ordered = np.sort(np.asarray(sample))
    if len(ordered) == 0:
        msg = "ccdf_at needs at least one value"
        raise ValueError(msg)
    below = np.searchsorted(ordered, marks, side="left")
    return (len(ordered) - below) / len(ordered)

local_slopes(sample, marks)

Compute the slope of the log-log CCDF between successive marks.

Parameters:

Name Type Description Default
sample array_like

The values; at least one.

required
marks sequence of int

Increasing marks; a slope is taken between each pair of neighbors.

required

Returns:

Type Description
DataFrame

Columns lower, upper, and slope, one row per pair of successive marks; slope is NaN where no value reaches the upper mark.

Examples:

A sample whose CCDF halves at each doubling has slope -1:

>>> slopes = local_slopes([1, 1, 2, 4], [1, 2, 4])
>>> [round(float(s), 3) for s in slopes["slope"]]
[-1.0, -1.0]
Source code in champalimaud/ccdf.py
def local_slopes(sample: ArrayLike, marks: Sequence[int]) -> pl.DataFrame:
    """Compute the slope of the log-log CCDF between successive marks.

    Parameters
    ----------
    sample : array_like
        The values; at least one.
    marks : sequence of int
        Increasing marks; a slope is taken between each pair of
        neighbors.

    Returns
    -------
    polars.DataFrame
        Columns ``lower``, ``upper``, and ``slope``, one row per pair
        of successive marks; ``slope`` is NaN where no value reaches
        the upper mark.

    Examples
    --------
    A sample whose CCDF halves at each doubling has slope -1:

    >>> slopes = local_slopes([1, 1, 2, 4], [1, 2, 4])
    >>> [round(float(s), 3) for s in slopes["slope"]]
    [-1.0, -1.0]
    """
    fraction = ccdf_at(sample, marks)
    lower = np.array(marks[:-1], dtype=float)
    upper = np.array(marks[1:], dtype=float)
    with np.errstate(divide="ignore", invalid="ignore"):
        slope = np.log(fraction[1:] / fraction[:-1]) / np.log(upper / lower)
    slope[~np.isfinite(slope)] = np.nan
    return pl.DataFrame({"lower": lower, "upper": upper, "slope": slope})