load¶
champalimaud.load
¶
The data files the project reads, and the readers for them.
The registry DATASETS lists every file or folder that fetch
writes.
status reports which are present and usable.
The readers are pure: no network and no writes, except that
figure_path creates config.FIGURES.
Tables come back as polars frames.
A missing file raises with the command that writes it.
Dataset
dataclass
¶
One data file or folder.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Short name, as shown in the status report. |
source |
str
|
Where it comes from, such as |
path |
Path
|
The file or folder, inside |
fetch |
str
|
The command that writes it. |
check |
callable
|
Takes |
partial |
(Path, optional)
|
A folder of pieces left by an interrupted fetch. |
Source code in champalimaud/load.py
check_csv(path, columns)
¶
Check that a CSV file opens and has some columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The file, gzipped or not. |
required |
columns
|
sequence of str
|
Columns that must be present. |
required |
Returns:
| Type | Description |
|---|---|
str or None
|
|
Examples:
>>> import tempfile
>>> folder = Path(tempfile.mkdtemp())
>>> _ = (folder / "a.csv").write_text("id,name\n1,x\n")
>>> check_csv(folder / "a.csv", ["id"])
>>> check_csv(folder / "a.csv", ["size"])
'missing columns: size'
Source code in champalimaud/load.py
check_parquet(path, columns)
¶
Check that a parquet file opens and has some columns.
Only the footer is read, so the check is fast on a large file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The file. |
required |
columns
|
sequence of str
|
Columns that must be present. |
required |
Returns:
| Type | Description |
|---|---|
str or None
|
|
Examples:
>>> import tempfile
>>> path = Path(tempfile.mkdtemp()) / "a.parquet"
>>> pl.DataFrame({"count": [1]}).write_parquet(path)
>>> check_parquet(path, ["count"])
>>> check_parquet(path, ["id"])
'missing columns: id'
Source code in champalimaud/load.py
check_swc_folder(path)
¶
Check that a folder holds skeletons and no unfinished download.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
The folder of |
required |
Returns:
| Type | Description |
|---|---|
str or None
|
|
Examples:
>>> import tempfile
>>> folder = Path(tempfile.mkdtemp())
>>> check_swc_folder(folder)
'no skeletons'
>>> _ = (folder / "1.swc").write_text("1 1 0 0 0 1 -1\n")
>>> check_swc_folder(folder)
Source code in champalimaud/load.py
figure_path(name)
¶
Return the path of a figure file, creating the folder.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
name
|
str
|
File name under |
required |
Returns:
| Type | Description |
|---|---|
Path
|
|
Source code in champalimaud/load.py
load_census()
¶
Read the cell-type census.
Only root_id and primary_type are read; the other columns
of the file, such as additional_type(s), are left out.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Source code in champalimaud/load.py
load_columns()
¶
Read the column assignment of the columnar cells.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Source code in champalimaud/load.py
load_connections_of(root_ids, *, end='pre', min_weight=1)
¶
Load the stored connections that start or end at some cells.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root_ids
|
sequence of int
|
The cells to select. |
required |
end
|
(pre, post)
|
Keep connections that start at ( |
"pre"
|
min_weight
|
int
|
Smallest weight, in synapses, to keep. |
1
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Raises:
| Type | Description |
|---|---|
ValueError
|
For any other |
Source code in champalimaud/load.py
load_example_synapses()
¶
Read every synapse of the example cell.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Source code in champalimaud/load.py
load_proofread_ids()
¶
Read the list of proofread neurons.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Column |
Source code in champalimaud/load.py
load_skeletons(census, types)
¶
Read the stored skeletons of some types.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
census
|
DataFrame
|
Has columns |
required |
types
|
sequence of str
|
The |
required |
Returns:
| Type | Description |
|---|---|
list of navis.TreeNeuron
|
One neuron per stored skeleton of those types, named
|
Source code in champalimaud/load.py
load_visual_types()
¶
Read the visual neuron types, with the side of each cell.
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Source code in champalimaud/load.py
require(path)
¶
Return a data path, or stop with the command that writes it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
Path
|
A data file or folder of |
required |
Returns:
| Type | Description |
|---|---|
Path
|
|
Raises:
| Type | Description |
|---|---|
SystemExit
|
When |
Source code in champalimaud/load.py
root_ids_of_type(census, cell_type)
¶
List the root ids of one primary type, as Python ints.
The ids are plain Python ints, not numpy integers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
census
|
DataFrame
|
Has columns |
required |
cell_type
|
str
|
The |
required |
Returns:
| Type | Description |
|---|---|
list of int
|
The |
Examples:
>>> import polars as pl
>>> census = pl.DataFrame(
... {
... "root_id": [1, 2, 3],
... "primary_type": ["A", "B", "A"],
... }
... )
>>> root_ids_of_type(census, "A")
[1, 3]
Source code in champalimaud/load.py
scan_connections()
¶
Scan every stored connection lazily.
Filter before calling collect on the result: the file holds
every pair of cells with a proofread end.
Returns:
| Type | Description |
|---|---|
LazyFrame
|
Columns |
Source code in champalimaud/load.py
status(datasets=DATASETS)
¶
Report whether each dataset is present and usable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
datasets
|
sequence of Dataset
|
What to check. |
`DATASETS`
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
One row per dataset, with columns:
|
Examples:
>>> import tempfile
>>> folder = Path(tempfile.mkdtemp())
>>> there = Dataset(
... "a", "x", folder, "f()", lambda p: None
... )
>>> gone = Dataset("b", "x", folder / "gone", "g()", lambda p: None)
>>> status([there, gone]).select("dataset", "state").rows()
[('a', 'ok'), ('b', 'missing')]
Source code in champalimaud/load.py
220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 | |