fetch¶
champalimaud.fetch
¶
Every network fetch in the project.
Writes the config data files that champalimaud.load reads.
FETCHERS maps each dataset of champalimaud.load.DATASETS to the
function that writes it.
The fetches need a saved CAVE token, and the census needs
CODEX_API_TOKEN if it is missing.
- cell types and other tables from Codex
- the connections of every proofread neuron, from CAVE
- skeletons, one SWC file per cell, from Codex
- every synapse of one example cell, from CAVE
FETCHERS = {'census': lambda: download_codex(Sources.CODEX_CENSUS, config.CENSUS), 'visual_types': lambda: download_codex('visual_neuron_types'), 'columns': lambda: download_codex('column_assignment'), 'classification': lambda: download_codex('classification'), 'connections': fetch_connections, 'skeletons': fetch_skeletons, 'example_synapses': fetch_example_synapses}
module-attribute
¶
The function that writes each dataset, by the name in
champalimaud.load.DATASETS, in the order to run them.
Inventory
¶
Sources
¶
Remote services and the ids of the products fetched from them.
Every id, URL, and limit below is a choice; the comment on each line says why.
Source code in champalimaud/fetch.py
TimeoutAdapter
¶
Bases: HTTPAdapter
Gives every request a timeout.
requests has none by default, and a connection that dies mid-query (a laptop suspended during a long fetch) then waits forever.
Source code in champalimaud/fetch.py
send(request, *args, **kwargs)
¶
Send a request, with a default timeout.
batch_connections(outputs, inputs, proofread)
¶
Count the pair weights of one batch of cells.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
outputs
|
DataFrame
|
Synapses whose presynaptic cell is in the batch. |
required |
inputs
|
DataFrame
|
Synapses whose postsynaptic cell is in the batch. Those from a proofread cell are dropped, since that cell's own batch counts them. |
required |
proofread
|
Series
|
Root ids of every proofread neuron. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Examples:
>>> import polars as pl
>>> outputs = pl.DataFrame(
... {"pre_pt_root_id": [2, 2, 2], "post_pt_root_id": [1, 9, 2]}
... )
>>> inputs = pl.DataFrame(
... {"pre_pt_root_id": [1, 9], "post_pt_root_id": [2, 2]}
... )
>>> batch_connections(outputs, inputs, pl.Series([1, 2])).rows()
[(2, 1, 1), (2, 9, 1), (9, 2, 1)]
Source code in champalimaud/fetch.py
cave_materialize()
¶
Create the CAVE materialize client for Sources.STACK.
Every request carries a timeout, from TimeoutAdapter.
Needs a saved CAVE token.
Returns:
| Type | Description |
|---|---|
MaterializationClient
|
The client. |
Source code in champalimaud/fetch.py
check_inventory(m)
¶
Read the materialization versions and the table names.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
m
|
MaterializationClient
|
As |
required |
Returns:
| Type | Description |
|---|---|
Inventory
|
The versions the server lists, the version in use, and the table names. |
Examples:
>>> class Stub:
... version = 783
... def get_versions(self):
... return [783]
... def get_tables(self):
... return ["nuclei_v1"]
>>> check_inventory(Stub())
Inventory(versions=[783], version=783, tables=['nuclei_v1'])
Source code in champalimaud/fetch.py
codex_url(product, api_token)
¶
Build the Codex download URL for one product.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
product
|
str
|
One of |
required |
api_token
|
str
|
The Codex account token. |
required |
Returns:
| Type | Description |
|---|---|
str
|
The URL, in |
Examples:
>>> codex_url("classification", "<token>")
'https://codex.flywire.ai/api/download_resource?data_product=classification&dataset=fafb&api_token=<token>'
Source code in champalimaud/fetch.py
download_codex(product, dest=None)
¶
Write one Codex product to a file.
Needs Sources.CODEX_ENV in the environment.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
product
|
str
|
One of |
required |
dest
|
Path
|
Where to write; by default |
None
|
Returns:
| Type | Description |
|---|---|
Path
|
The file written. |
Raises:
| Type | Description |
|---|---|
SystemExit
|
When |
Source code in champalimaud/fetch.py
drop_self_edges(syn)
¶
Remove synapses whose pre and post root ids are the same cell.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
syn
|
DataFrame
|
Has |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
The rows of |
Examples:
>>> import pandas as pd
>>> syn = pd.DataFrame(
... {"pre_pt_root_id": [1, 2], "post_pt_root_id": [1, 3]}
... )
>>> drop_self_edges(syn)["pre_pt_root_id"].tolist()
[2]
Source code in champalimaud/fetch.py
fetch_connection_shard(ids, proofread, dest)
¶
Fetch the connections of some cells and write their pair weights.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ids
|
list of int
|
The cells of this batch. |
required |
proofread
|
Series
|
Root ids of every proofread neuron. |
required |
dest
|
Path
|
Parquet file to write; it is written under a temporary name first, so a fetch that dies halfway leaves no truncated file. |
required |
Returns:
| Type | Description |
|---|---|
int
|
The number of pairs written. |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
From |
Source code in champalimaud/fetch.py
458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 | |
fetch_connections(workers=4)
¶
Fetch every proofread cell's connections and merge them.
The connections go in batches of Sources.BATCH cells to
config.CONNECTIONS_PARTS, and a rerun skips the batches that
are there; the merge writes config.CONNECTIONS and deletes the
folder.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
workers
|
int
|
Batches fetched at once. |
4
|
Raises:
| Type | Description |
|---|---|
SystemExit
|
When any batch failed; rerun to fetch the missing ones. |
Source code in champalimaud/fetch.py
536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 | |
fetch_example_synapses(cell_type=EXAMPLE_TYPE)
¶
Write every synapse of the first cell of a type.
The rows are those of Sources.SYNAPSE_VIEW, the table that
fetch_connections counts, with the synapses of the cell onto
itself dropped.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cell_type
|
str
|
The |
`EXAMPLE_TYPE`
|
Returns:
| Type | Description |
|---|---|
Path
|
|
Source code in champalimaud/fetch.py
fetch_skeletons(types=SKELETON_TYPES)
¶
Write the first skeletons of each type to config.SKELETONS.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
types
|
sequence of str
|
The |
`SKELETON_TYPES`
|
Returns:
| Type | Description |
|---|---|
list of pathlib.Path
|
The |
Source code in champalimaud/fetch.py
ids_in_nuclei(m, root_ids)
¶
Count the root ids that nuclei_v1 has a row for.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
m
|
MaterializationClient
|
As |
required |
root_ids
|
sequence of int
|
The ids to look up. |
required |
Returns:
| Type | Description |
|---|---|
int
|
How many of them are in the table. |
Examples:
>>> import pandas as pd
>>> class Stub:
... def query_table(self, table, **kwargs):
... return pd.DataFrame({"pt_root_id": [1, 1, 2]})
>>> ids_in_nuclei(Stub(), [1, 2, 3])
2
Source code in champalimaud/fetch.py
query_synapses(m, column, ids)
¶
Query the synapse rows with a column in some root ids.
A query that hits Sources.ROW_CAP is truncated without an error,
so it is split in half until each answer is under the cap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
m
|
MaterializationClient
|
As |
required |
column
|
(pre_pt_root_id, post_pt_root_id)
|
The column to test. |
"pre_pt_root_id"
|
ids
|
list of int
|
Root ids to look for in |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Columns |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
When one root id alone has at least |
Source code in champalimaud/fetch.py
sample_tables(m)
¶
Query a few rows of the tables every fetch reads.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
m
|
MaterializationClient
|
As |
required |
Returns:
| Type | Description |
|---|---|
nuclei, view_row, synapse : pandas.DataFrame
|
Three rows of |
Examples:
>>> import pandas as pd
>>> class Stub:
... def query_table(self, table, **kwargs):
... return pd.DataFrame({"table": [table]})
... def query_view(self, view, **kwargs):
... return pd.DataFrame({"table": [view]})
>>> nuclei, view_row, synapse = sample_tables(Stub())
>>> nuclei["table"].tolist()
['nuclei_v1']
Source code in champalimaud/fetch.py
skeleton_url(root_id)
¶
Build the Codex SWC URL for one root id.
This is the only skeleton source besides the bulk zip on their download page.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root_id
|
int
|
The cell. |
required |
Returns:
| Type | Description |
|---|---|
str
|
The URL under |
Examples: