Skip to content

Wizard

Dataset format

What a task model's training data looks like, and complete scripts that turn an EDF file, a BIDS dataset or a NumPy array into it.

The format

A task model, an instance of Wizard, the mixture model, learns from trials: one epoch of EEG per event, each with one label. Whatever your data starts as, it goes to the API as:

What to send
Trialsx shaped (trials, channels, samples): one row per channel, the same length for every trial. As nested lists, or as base64 little-endian float32 with shape (smaller and faster).
Labelslabels: one string per trial, in the order of x. At least two classes; at least 5 trials per class are required, and 20 or more are recommended.
Channelschannels: one 10-10 or 10-20 name per row, any layout (3, 19, 64 channels…). Predictions must send the same names, in any order.
Unitsunits: uV, mV or V, the unit x is actually in. Normalized or unitless data is refused.
Sampling ratesampling_rate_hz: the rate x is actually at. The API resamples; it cannot detect a wrong rate.
Sessionssessions (optional): one session name per trial. With two or more sessions, validation keeps sessions apart, so the scores say how the model does on a session it has not seen.
Descriptiondescription: at most 100 characters on what the task is and what each label means.

Channel names

Every channel must carry its 10-10 name (the 10-20 names are among them): Fp1, FC3, Cz, POz. The API matches names, never positions in a file, so:

  • Clean file-style labels. EDF files often say EEG Fp1-REF or Fp1.; send Fp1. The API does not strip prefixes or reference suffixes.
  • Legacy temporal names are fine. T3, T4, T5 and T6 are accepted and read as T7, T8, P7 and P8, the same electrodes. The helper below renames them anyway, for tidy channel lists.
  • Re-reference a bipolar montage. A channel such as Fp1-F7 is the difference of two electrodes, not one electrode against a common reference. Re-reference it before sending; do not relabel it with its first electrode.
  • Drop what is not EEG. EOG, ECG, EMG, trigger and status channels, and the reference electrodes themselves, have no 10-10 name to send.

BioSemi-style labels (A1 … B32) are positions on a particular cap, not 10-10 names: map them with your cap’s layout first.

Trial length

Epoch each trial from its event marker; the API does not see your triggers. Short trials are fine: a method whose window would be more than half zero padding is not tried when another method fits the trial, so 1-second ERP epochs go to methods with short windows.

TaskEpoch per trialIn the scripts below
ERP (event-related potentials)1 s from the stimulus (0 to 1 s)START, SECONDS = 0.0, 1.0
Motor imagery or movement0 to 4 s from the cue (1 to 4 s also works: it leaves out the response to the cue itself)START, SECONDS = 0.0, 4.0
Sleep staging30 s scoring epochsSTART, SECONDS = 0.0, 30.0, chunk=30

For sleep, chunk=30 cuts each long sleep-stage annotation into back-to-back 30-second trials, so a 10-minute “Sleep stage 2” annotation gives twenty trials labeled N2 (or whatever you map it to).

Size limits and batching

One request carries at most 4,000,000 numbers (trials × channels × samples), counted as sent and again after resampling to 256 Hz, at most 4,096 trials, and a body of at most 64 MiB. At 256 Hz a 4-second trial of 19 channels is 19,456 numbers, so a request holds about 200 of them; a 30-second sleep trial of 19 channels is 145,920 numbers, about 27 per request.

For more trials than one request holds, create the model with "defer_fit": true and the first batch, add the rest with POST /v1/task-models/{id}/trials, then call POST /v1/task-models/{id}/fit within 7 days: a collecting model that is never fitted is deleted with its collected features. Mix the classes across batches: the helper below shuffles the trials once, so every batch carries every class.

Tutorial

Three complete scripts take you from a file to a task model: (a) an EDF or BDF recording with annotations, (b) a BIDS-EEG dataset, (c) a NumPy array you already have. They share one helper, and (d) predicts on a new recording. Install the dependencies in a virtual environment (see the Quickstart), with mne-bids only for (b):

Terminal
python -m pip install mne numpy requests mne-bids

Every script starts with DRY_RUN = True: the first run asks for an estimate (the create fee and any warnings) and charges nothing. Read them, set DRY_RUN = False and run it again to create the model. These scripts are Python only, because MNE and mne-bids have no JavaScript counterpart; the requests they make are shown in Python, JavaScript and cURL on Wizard.

The helper: dydema_dataset.py

Save this as dydema_dataset.py next to your scripts. It:

  • cleans channel labels (EEG Fp1-REF becomes Fp1, and, for tidiness, T3 becomes T7), keeps every EEG channel with a 10-10 name and drops the rest (EOG, ECG, triggers, channels marked bad), and stops on a bipolar pair;
  • resamples to 256 Hz after dropping, and cuts one trial per annotation, in microvolts;
  • checks the classes, shuffles the trials once, and uploads them in batches under the size limits: one create if they fit, else create with defer_fit, /trials for the rest, and /fit;
  • saves its progress in task_model.json after every request, with the channel names, so a run refused halfway (out of credit, say) continues where it stopped, and predictions later send the same channels.
dydema_dataset.py
Python
"""dydema_dataset.py: EEG recordings -> task-model trials, and the requests that upload them.

Save this file next to your scripts; they import it. Needs: python -m pip install mne numpy requests
"""
import base64
import json
import os
from collections import Counter

import mne
import numpy as np
import requests

API = 'https://dydema--eegapi-gateway.modal.run/v1'
RATE = 256  # the scripts resample to 256 Hz before sending, so the size of every request is predictable
MAX_VALUES = 4_000_000  # numbers per request (trials x channels x samples), as sent and after resampling to 256 Hz
MAX_TRIALS = 4_096  # trials (and four-second windows) per request
MIN_PER_CLASS, GOOD_PER_CLASS = 5, 20  # trials per class: the API's minimum, and what we recommend

# The 10-10 electrode names (the 10-20 names are among them). A channel whose cleaned label is not one of these
# (EOG, ECG, EMG, a trigger or status channel, a reference electrode) is not EEG the model can place: it is dropped.
TEN_TEN = """Fp1 Fpz Fp2 AF9 AF7 AF5 AF3 AF1 AFz AF2 AF4 AF6 AF8 AF10 F9 F7 F5 F3 F1 Fz F2 F4 F6 F8 F10
FT9 FT7 FC5 FC3 FC1 FCz FC2 FC4 FC6 FT8 FT10 T9 T7 C5 C3 C1 Cz C2 C4 C6 T8 T10
TP9 TP7 CP5 CP3 CP1 CPz CP2 CP4 CP6 TP8 TP10 P9 P7 P5 P3 P1 Pz P2 P4 P6 P8 P10
PO9 PO7 PO5 PO3 PO1 POz PO2 PO4 PO6 PO8 PO10 O9 O1 Oz O2 O10 I1 Iz I2""".split()
LEGACY = {"T3": "T7", "T4": "T8", "T5": "P7", "T6": "P8"}  # older names: the API accepts both; renamed for tidiness
REFERENCES = {"REF", "LE", "A1", "A2", "M1", "M2", "AVG"}  # reference suffixes, as in "EEG Fp1-REF"
TO_MICROVOLTS = {"uV": 1.0, "mV": 1e3, "V": 1e6}


def electrode(label):
    """'EEG Fp1-REF' -> 'Fp1', 'T3' -> 'T7', 'EOG left' -> None. Stops on a bipolar pair such as 'Fp1-F7'."""
    name = label.strip().rstrip(".")  # PhysioNet writes "Fp1." and "T7.."
    if name.upper().startswith("EEG "):
        name = name[4:]
    name, _, suffix = name.partition("-")
    name = LEGACY.get(name.strip().upper(), name.strip())
    match = next((c for c in TEN_TEN if c.lower() == name.lower()), None)
    if match and suffix and suffix.strip().upper() not in REFERENCES:
        raise SystemExit(
            f"channel {label!r} is a bipolar pair (bipolar montage): re-reference to a common reference first. "
            f"If {suffix!r} is your reference electrode, add it to REFERENCES."
        )
    return match


def clean_names(labels):
    """{file label: 10-10 name} for every EEG channel with a 10-10 name, in file order; the rest are left out."""
    names = {}
    for label in labels:
        name = electrode(label)
        if name is None:
            continue
        if name in names.values():
            raise SystemExit(f"two channels clean to {name!r}: keep one of them")
        names[label] = name
    if not names:
        raise SystemExit(
            f"no channel has a 10-10 name (labels: {labels[:8]}...). Rename them first: BioSemi-style labels "
            "such as 'A1'...'B32' need your cap's layout to become 10-10 names."
        )
    return names


def prepare(raw):
    """Clean labels, keep the EEG channels with 10-10 names (not those marked bad), resample to 256 Hz."""
    names = {label: name for label, name in clean_names(raw.ch_names).items() if label not in raw.info["bads"]}
    dropped = [label for label in raw.ch_names if label not in names]
    print(f"keeping {len(names)} EEG channels; dropping {len(dropped)}: {dropped}")
    raw.pick(list(names)).load_data()  # keep the EEG channels, THEN resample: resampling what you drop is waste
    raw.rename_channels(names)
    raw.set_channel_types({name: "eeg" for name in names.values()})  # so get_data(units="uV") converts every row
    raw.resample(RATE)
    return raw


def read_recording(path):
    """An EDF, BDF (or any file MNE reads) as a prepared recording: 10-10 channels at 256 Hz, annotations kept."""
    return prepare(mne.io.read_raw(path, preload=False, verbose="error"))


def trials_from_annotations(raw, cues, start=0.0, seconds=4.0, chunk=None):
    """One trial per annotation named in cues ({annotation: label}), start s after it, seconds long.

    chunk cuts long annotations into back-to-back pieces first (sleep: chunk=30 turns a 10-minute "Sleep stage 2"
    annotation into twenty 30-second trials). Returns (x, labels): x in microvolts, (trials, channels, samples).
    """
    events, ids = mne.events_from_annotations(raw, chunk_duration=chunk, verbose="error")
    print("annotations in the recording:", sorted(ids))
    names = {code: description for description, code in ids.items()}
    rate, data = raw.info["sfreq"], raw.get_data(units="uV")  # MNE stores volts; the scripts send microvolts
    length = round(seconds * rate)
    x, labels = [], []
    for sample, _, code in events:
        if names[code] not in cues:
            continue
        first = sample - raw.first_samp + round(start * rate)
        if first < 0 or first + length > data.shape[1]:
            print(f"skipped a {names[code]!r} trial at {(sample - raw.first_samp) / rate:.1f} s: it runs past the recording")
            continue
        x.append(data[:, first:first + length])
        labels.append(cues[names[code]])
    if not x:
        raise SystemExit(f"no trials: none of {sorted(cues)} is among the annotations printed above")
    return np.stack(x).astype("float32"), labels


def from_array(x, channels, units="uV"):
    """An array you already have, (trials, channels, samples): cleaned channel names, non-EEG rows dropped, microvolts."""
    names = clean_names(list(channels))
    rows = [i for i, label in enumerate(channels) if label in names]
    return np.asarray(x, dtype="float32")[:, rows] * TO_MICROVOLTS[units], [names[channels[i]] for i in rows]


def stack_sessions(parts):
    """[(x, labels, channels, session), ...] -> one dataset on the channels every session has, with a session per trial."""
    common = [c for c in parts[0][2] if all(c in p[2] for p in parts)]
    if len(common) < len(parts[0][2]):
        print(f"keeping the {len(common)} channels every session has")
    x = np.concatenate([p[0][:, [p[2].index(c) for c in common]] for p in parts])
    labels = [label for p in parts for label in p[1]]
    sessions = [p[3] for p in parts for _ in p[1]]
    return x, labels, sessions, common


def check_classes(labels):
    counts = Counter(labels)
    print(f"{len(labels)} trials: {dict(counts)}")
    if len(counts) < 2 or min(counts.values()) < MIN_PER_CLASS:
        raise SystemExit(f"a task model needs at least 2 classes of at least {MIN_PER_CLASS} trials each")
    if min(counts.values()) < GOOD_PER_CLASS:
        print(f"note: fewer than {GOOD_PER_CLASS} trials in a class; the model will be noisy")


def batch_size(x, rate):
    """Trials per request: under the value cap as sent and after resampling to 256 Hz, and under 4,096 windows."""
    values = x.shape[1] * x.shape[2] * max(1.0, 256 / rate)
    windows = max(1, int(x.shape[2] / rate // 4))  # a trial of 8 s or more is several four-second windows
    return max(1, min(int(MAX_VALUES // values), MAX_TRIALS // windows))


def post(path, body):
    response = requests.post(
        f"{API}{path}", headers={"Authorization": f"Bearer {os.environ['DYDEMA_API_KEY']}"}, json=body, timeout=300
    )
    if not response.ok:
        error = response.json()
        raise SystemExit(f"{response.status_code} {error['code']}: {error['message']} {error['alternatives']}")
    return response.json()


def signal(x, channels, rate):
    x = np.ascontiguousarray(x, dtype="<f4")  # little-endian float32, sent as base64 with its shape
    return {"x": base64.b64encode(x.tobytes()).decode("ascii"), "shape": list(x.shape), "channels": list(channels),
            "sampling_rate_hz": rate, "units": "uV"}


def create_task_model(x, labels, channels, rate, description, name=None, sessions=None,
                      expires_in_days=None, dry_run=True, state="task_model.json"):
    """Create a task model from every trial: one request if they fit, else create (defer_fit) + /trials + /fit.

    dry_run=True (the default) only asks for the estimated cost and limitations: nothing is fitted or
    charged. Read it, then call again with dry_run=False. Progress is saved in state after every request, so
    running again after a refusal (402 out of credit, say) continues from the next batch. Delete it to start over.
    """
    check_classes(labels)
    order = np.random.default_rng(0).permutation(len(labels))  # shuffled once, so every batch holds every class
    x, labels = x[order], [labels[i] for i in order]
    sessions = None if sessions is None else [sessions[i] for i in order]
    per = batch_size(x, rate)

    def batch(first):
        body = signal(x[first:first + per], channels, rate) | {"labels": labels[first:first + per]}
        return body | ({"sessions": sessions[first:first + per]} if sessions is not None else {})

    def save(task, sent):  # the channels too: predictions must send the same channel names
        with open(state, "w") as f:
            json.dump({"id": task["id"], "sent": sent, "channels": list(channels), "task_model": task}, f, indent=2)

    first_request = batch(0) | {"description": description, "defer_fit": per < len(x)}
    first_request |= {k: v for k, v in {"name": name, "expires_in_days": expires_in_days}.items() if v is not None}
    if dry_run:
        preview = post("/task-models", first_request | {"dry_run": True})
        batches = -(-len(x) // per)
        print(f"dry run ({batches} request(s) of up to {per} trials; this one carried the first batch):")
        print(json.dumps(preview, indent=2))
        return preview
    if os.path.exists(state):
        with open(state) as f:
            saved = json.load(f)
        task, sent = saved["task_model"], saved["sent"]
        print(f"continuing task model {task['id']} from {state}: {sent} of {len(x)} trials already sent")
    else:
        task, sent = post("/task-models", first_request), min(per, len(x))
        save(task, sent)
        print(f"created {task['id']} ({task['status']})")
    for first in range(sent, len(x), per):
        post(f"/task-models/{task['id']}/trials", batch(first))
        save(task, min(first + per, len(x)))
        print(f"sent {min(first + per, len(x))} of {len(x)} trials")
    if task["status"] == "collecting":
        task = post(f"/task-models/{task['id']}/fit", {})
        save(task, len(x))
    print(f"task model {task['id']} is {task['status']}; validation results:")
    print(json.dumps(task["training_report"], indent=2))
    return task


def predict(task_model_id, x, channels, rate):
    """One prediction per trial ({"label", "probabilities"}, in order). Send the channel names the model was fitted on."""
    per, out = batch_size(x, rate), []
    for first in range(0, len(x), per):
        out += post(f"/task-models/{task_model_id}/predictions", signal(x[first:first + per], channels, rate))["data"]
    return out

(a) EDF or BDF with annotations

Set PATH to your file, CUES to the annotations that mark trials and the labels you want for them, and DESCRIPTION to your task. The script prints the annotations it finds, so a first run tells you what to put in CUES.

edf_to_task_model.py
Python
"""An EDF or BDF recording with trial annotations -> a task model. Run: python edf_to_task_model.py"""
from dydema_dataset import RATE, create_task_model, read_recording, trials_from_annotations

PATH = "recording.edf"  # your file: EDF or BDF, with an annotation at every cue
CUES = {"T1": "left", "T2": "right"}  # annotation in the file -> the label you want (the API reads the labels)
START, SECONDS = 0.0, 4.0  # motor imagery: 0-4 s after the cue. ERP: 0.0, 1.0. Sleep: 0.0, 30.0 with chunk=30
DESCRIPTION = (
    "Motor imagery: after a cue the participant imagines squeezing the left or the right hand. "
    "Labels: left = left-hand imagery, right = right-hand imagery."
)
DRY_RUN = True  # True: the estimated cost and limitations, nothing charged. Then set False to create.

raw = read_recording(PATH)  # labels cleaned, EEG channels with 10-10 names kept, the rest dropped, 256 Hz
x, labels = trials_from_annotations(raw, CUES, start=START, seconds=SECONDS)
print("trials:", x.shape, "(trials, channels, samples) in microvolts")
task = create_task_model(x, labels, raw.ch_names, RATE, DESCRIPTION, name="hand imagery",
                         expires_in_days=90, dry_run=DRY_RUN)

(b) A BIDS-EEG dataset

mne-bids reads each recording with its *_events.tsv (the trial_type column becomes the annotations) and its *_channels.tsv (channels marked bad are dropped). Each BIDS session becomes the sessions entry of its trials, and only the channels every session has are kept.

bids_to_task_model.py
Python
"""A BIDS-EEG dataset (events in *_events.tsv) -> a task model. Needs: python -m pip install mne-bids"""
from mne_bids import BIDSPath, get_entity_vals, read_raw_bids
from dydema_dataset import RATE, create_task_model, prepare, stack_sessions, trials_from_annotations

ROOT, SUBJECT, TASK = "bids_dataset", "01", "handimagery"  # your dataset's folder, subject and task
CUES = {"left_hand": "left", "right_hand": "right"}  # trial_type in events.tsv -> the label you want
START, SECONDS = 0.0, 4.0  # motor imagery: 0-4 s after the cue. ERP: 0.0, 1.0. Sleep: 0.0, 30.0 with chunk=30
DESCRIPTION = (
    "Motor imagery: after a cue the participant imagines squeezing the left or the right hand. "
    "Labels: left = left-hand imagery, right = right-hand imagery."
)
DRY_RUN = True  # True: the estimated cost and limitations, nothing charged. Then set False to create.

parts = []
for session in get_entity_vals(ROOT, "session") or [None]:  # a dataset without sessions is one session
    path = BIDSPath(root=ROOT, subject=SUBJECT, session=session, task=TASK, datatype="eeg")
    raw = prepare(read_raw_bids(path, verbose="error"))  # events.tsv becomes annotations; bad channels are dropped
    x, labels = trials_from_annotations(raw, CUES, start=START, seconds=SECONDS)
    parts.append((x, labels, raw.ch_names, session or "1"))
    print(f"session {session}: {len(labels)} trials")

x, labels, sessions, channels = stack_sessions(parts)  # one session per trial: validation across sessions
task = create_task_model(x, labels, channels, RATE, DESCRIPTION, name="hand imagery (BIDS)",
                         sessions=sessions, expires_in_days=90, dry_run=DRY_RUN)

(c) A NumPy array you already have

Already epoched? Save x, labels and channels in an .npz file, or change the loading line. The helper cleans the channel names and drops the non-EEG rows; declare the rate and units your array is actually in.

numpy_to_task_model.py
Python
"""Trials you already have as a NumPy array -> a task model. Run: python numpy_to_task_model.py"""
import numpy as np
from dydema_dataset import create_task_model, from_array

data = np.load("trials.npz")  # yours: x (trials, channels, samples), labels (one per trial), channels (one per row)
x, channels = from_array(data["x"], [str(c) for c in data["channels"]], units="uV")  # "uV", "mV" or "V"
labels = [str(label) for label in data["labels"]]
RATE = 250  # the rate x was recorded at: the API resamples, so there is no need to resample it yourself
DESCRIPTION = (
    "Auditory oddball: rare 1000 Hz tones among frequent 500 Hz tones, trials 0-1 s after tone onset. "
    "Labels: target = rare tone, standard = frequent tone."
)
DRY_RUN = True  # True: the estimated cost and limitations, nothing charged. Then set False to create.

task = create_task_model(x, labels, channels, RATE, DESCRIPTION, name="oddball",
                         expires_in_days=30, dry_run=DRY_RUN)

(d) Predict on a new recording

Reads the id and the channel names from task_model.json, keeps exactly those channels of the new recording, and prints one predicted label per trial.

predict_recording.py
Python
"""New trials from another recording -> one predicted label per trial, from the task model in task_model.json."""
import json
from collections import Counter
from dydema_dataset import RATE, predict, read_recording, trials_from_annotations

with open("task_model.json") as f:  # written by create_task_model
    saved = json.load(f)
fitted_on = saved["channels"]  # the channel names the model was fitted on; predictions must send the same names

raw = read_recording("new_recording.edf")  # your file
missing = [c for c in fitted_on if c not in raw.ch_names]
if missing:
    raise SystemExit(f"the new recording has no {missing}; the model was fitted on {fitted_on}")
raw.pick(fitted_on)  # exactly the fitted channels (the order does not matter to the API)
x, truth = trials_from_annotations(raw, {"T1": "left", "T2": "right"}, start=0.0, seconds=4.0)

predictions = predict(saved["id"], x, raw.ch_names, RATE)
predicted = [p["label"] for p in predictions]
print(Counter(predicted), "first trial:", predictions[0]["probabilities"])
print(f"accuracy against the file's own labels: {sum(p == t for p, t in zip(predicted, truth)) / len(truth):.2f}")
Check before you send: the units are the unit your file is really in (MNE reads EDF and BDF in volts and get_data(units="uV") converts), the trials are cut at the right place (print a few event times against your paradigm), and every class you expect is in the label counts the helper prints.
What a run prints (EDF example, dry run)
keeping 40 EEG channels; dropping 3: ['EOG1', 'ECG', 'Trigger']
annotations in the recording: ['T1', 'T2']
trials: (120, 40, 1024) (trials, channels, samples) in microvolts
120 trials: {'left': 60, 'right': 60}
dry run (2 request(s) of up to 97 trials; this one carried the first batch):
{ ...the category, the warnings and the estimated cost... }