Skip to content

overlap

champalimaud.overlap

Overlap between sets, as the Jaccard index.

The Jaccard index of two sets is the size of their intersection over the size of their union: 1 for the same set and 0 for disjoint sets.

jaccard(a, b)

Jaccard index of two sets, 0 for two empty sets.

Parameters:

Name Type Description Default
a set

The sets to compare.

required
b set

The sets to compare.

required

Returns:

Type Description
float

|a & b| / |a | b|.

Examples:

>>> jaccard({"A", "B", "C"}, {"B", "C", "D"})
0.5
>>> jaccard(set(), set())
0.0
Source code in champalimaud/overlap.py
def jaccard(a: set, b: set) -> float:
    """Jaccard index of two sets, 0 for two empty sets.

    Parameters
    ----------
    a, b : set
        The sets to compare.

    Returns
    -------
    float
        ``|a & b| / |a | b|``.

    Examples
    --------
    >>> jaccard({"A", "B", "C"}, {"B", "C", "D"})
    0.5
    >>> jaccard(set(), set())
    0.0
    """
    union = a | b
    if not union:
        return 0.0
    return len(a & b) / len(union)

pairwise_jaccard(sets)

The Jaccard index of every pair of the sets.

Parameters:

Name Type Description Default
sets dict of str to set

The sets, by name.

required

Returns:

Type Description
DataFrame

One row per pair of names, with columns a, b, jaccard, shared (the size of the intersection), and union (the size of the union).

See Also

jaccard : The index of one pair. sets_by_type : Builds the sets from a table.

Examples:

>>> table = pairwise_jaccard({"P": {1, 2}, "Q": {2, 3}})
>>> table.select("a", "b", "shared", "union").rows()
[('P', 'Q', 1, 3)]
>>> round(table["jaccard"].item(), 3)
0.333
Source code in champalimaud/overlap.py
def pairwise_jaccard(sets: dict[str, set]) -> pl.DataFrame:
    """The Jaccard index of every pair of the sets.

    Parameters
    ----------
    sets : dict of str to set
        The sets, by name.

    Returns
    -------
    polars.DataFrame
        One row per pair of names, with columns ``a``, ``b``,
        ``jaccard``, ``shared`` (the size of the intersection), and
        ``union`` (the size of the union).

    See Also
    --------
    jaccard : The index of one pair.
    sets_by_type : Builds the sets from a table.

    Examples
    --------
    >>> table = pairwise_jaccard({"P": {1, 2}, "Q": {2, 3}})
    >>> table.select("a", "b", "shared", "union").rows()
    [('P', 'Q', 1, 3)]
    >>> round(table["jaccard"].item(), 3)
    0.333
    """
    return pl.DataFrame(
        [
            {
                "a": a,
                "b": b,
                "jaccard": jaccard(sets[a], sets[b]),
                "shared": len(sets[a] & sets[b]),
                "union": len(sets[a] | sets[b]),
            }
            for a, b in combinations(sets, 2)
        ],
        schema={
            "a": pl.String,
            "b": pl.String,
            "jaccard": pl.Float64,
            "shared": pl.Int64,
            "union": pl.Int64,
        },
    )

sets_by_type(strong, column, types)

The values of a column among the rows of each type, as sets.

Parameters:

Name Type Description Default
strong DataFrame

Has a column type and the column named by column.

required
column str

Column whose values are collected.

required
types sequence of str

The types to collect, in order.

required

Returns:

Type Description
dict of str to set

One set per type; a type with no row has an empty set.

See Also

pairwise_jaccard : The overlap of these sets.

Examples:

>>> import polars as pl
>>> strong = pl.DataFrame(
...     {"type": ["P", "P", "Q"], "partner": [1, 2, 2]}
... )
>>> sets_by_type(strong, "partner", ["P", "Q", "R"])
{'P': {1, 2}, 'Q': {2}, 'R': set()}
Source code in champalimaud/overlap.py
def sets_by_type(
    strong: pl.DataFrame, column: str, types: Sequence[str]
) -> dict[str, set]:
    """The values of a column among the rows of each type, as sets.

    Parameters
    ----------
    strong : polars.DataFrame
        Has a column ``type`` and the column named by `column`.
    column : str
        Column whose values are collected.
    types : sequence of str
        The types to collect, in order.

    Returns
    -------
    dict of str to set
        One set per type; a type with no row has an empty set.

    See Also
    --------
    pairwise_jaccard : The overlap of these sets.

    Examples
    --------
    >>> import polars as pl
    >>> strong = pl.DataFrame(
    ...     {"type": ["P", "P", "Q"], "partner": [1, 2, 2]}
    ... )
    >>> sets_by_type(strong, "partner", ["P", "Q", "R"])
    {'P': {1, 2}, 'Q': {2}, 'R': set()}
    """
    return {
        t: set(values.to_list())
        for t, values in values_by_type(strong, column, types).items()
    }