Open research notebook 001—013

Signals,
before stories.

Project subtitle
An LLM’s interpretation of your brainwaves

EEGT studies whether a small vocabulary of numeric tokens can retain useful structure in EEG recordings. The current study compares waveform landmarks with numerical EEG embeddings; it runs no LLM decoder and does not interpret thoughts. The notebook shows the inputs, tests, and limits behind each result.

Token IDs describe measured signal patterns. They are not words, thoughts, diagnoses, or interpretations of a person.

FIG. 01 WAVEFORM → TOKEN
amplitudetime →
Schematic waveform divided into windows An invented line illustrates successive numeric signal windows. It is not recorded EEG or a result.
03110307
Schematic only. Invented shape and IDs show the process, not a recording, decoded meaning, or measured result.

Current protocolConditional timing control

InputOpen around-ear EEG

Release ruleNumbers after verification

01 / PURPOSE

Can a compact vocabulary keep the shape of a signal?

EEGT compares numerical signal codes and continuous representations, then measures their stability and limits across recordings. Prior exposure and training overlap are recorded separately for each experiment.

01

Observe

Begin with open EEG recordings and preserve the original samples. Numeric quality checks describe what can be analyzed.

02

Compress

Learn unsupervised codebooks from waveform, spectrum, and time-frequency views. Each token is an index in a learned vocabulary.

03

Challenge

Compare multiple seeds and vocabulary sizes, including on held-out recordings. Report numerical agreement and failure cases.

02 / INPUT CONTRACT

Raw numbers in. Claims stay bounded.

The discovery input is signal values and the numeric information needed to place them in time and channels. Task, identity, and clinical labels do not teach the codebook; metadata is kept separately for scientific audit.

Read the reproducible protocol
INPUT / OUTPUTTYPE CONTRACT · NO SAMPLE DATA
inputsample index × channel value
contextsampling rate + fixed processing
outputwindow → integer token ID

No words, task labels, identity labels, or diagnoses in discovery.

03 / ACQUISITION SCOPE

Compatible in format is not validated on hardware.

The Neurable Research Kit publicly advertises 12 EEG channels at 500 Hz. The open cEEGrid recordings below share an around-ear measurement geometry, but they were not captured with that headset. An import adapter cannot establish physical headset performance.

Sources and their role in this notebook
SourceMeasurementHere
Neurable Research Kit Advertised: 12 EEG channels, 500 HzTarget for a numeric import contract; no kit capture tested here.
OpenNeuro ds004015 Selected cEEGrid around-ear recordings, 500 HzOpen source recordings for Experiment 002.
OpenNeuro ds005207 Selected cEEGrid around-ear recordings, 250 HzOpen source recordings for Experiment 002.

Rates describe the selected source recordings. Source datasets can contain other acquisition types or rates; the run manifest must identify each selected file.

04 / RESEARCH NOTES

A ledger, not a reveal.

Each entry separates fixed inputs from measured outcomes and states what a result can support.

Latest entry first
2026-09-28 · Exploratory controlIndependent review accepted

What remains after event timing is shuffled?

Twelve existing blocks from six people give mean raw matching F1 of 0.8478 versus a mean timing-null median of 0.6634. The null preserves local event counts but breaks fine spacing and waveform constraints; the gap does not establish neural meaning.

6 people · 12 sessions · zero new source hoursRead the timing-control study
NOTE 013 / FIXED METHODS, NEW RECORDINGSFour estimable · all inconclusive

Keeping new people separate from new nights.

Twenty pinned recordings yielded 9,600 candidates and 174 fixed selections. Six complete new-session participants support four comparisons; adjusted p values are 1.00, 0.25, 0.50 and 1.00.

Only two new-person participants met the input rules, so that cohort remains unestimable. Independent recorded-ledger reproduction verified all 4,002 event partitions; neither this result nor model agreement establishes semantic meaning.

151.85 source hours · 87 selected minutesConsolidated evidence catalog · Read Experiment 013
NOTE 012 / DIRECT WAVEFORM LANDMARKSFour tests · all inconclusive

Measuring where the waveform turns.

We measured extrema, inflections, cycles and bursts on the same 122 selected segments, retaining every candidate and rejection across 2,806 event partitions.

Event-to-encoder alignment differs between CodeBrain and CBraMod. Adjusted p values are 0.50, 0.25, 1.00 and 1.00. Recurrent numerical turns do not yet establish universal tokens or transitions between brain states.

Five paired people · no new source hoursRead Experiment 012
NOTE 011 / TWO PRETRAINED ENCODERSPositive contrasts · tests inconclusive

Two encoders. Some alignment. An open question.

CodeBrain and CBraMod receive the same 122 selected wave segments and controls. Within-block original-minus-phase contrasts are positive in all five paired people.

Geometry and change contrasts have adjusted p = 0.125. Five-person test resolution, learned priors, shared recording effects and unknown training overlap keep the universal-geometry question open.

610 new passes · five paired peopleRead Experiment 011
NOTE 010 / TIME-DISTRIBUTED SAMPLEWeak geometry · transition test inconclusive

More usable waves. Weak shared geometry.

A frozen quality census checks 5,760 blocks across already examined recordings. It selects 122 time-distributed inputs; five people supply sufficient paired-night support, and the pretrained encoder completes 610 passes.

Geometry correlations are small, even where they exceed the specified shift controls. Change-profile tests do not clear correction. Every rejected and qualified-but-unselected block remains traceable.

Five paired people · no new source hoursRead Experiment 010
NOTE 009 / PRETRAINED ENCODERInsufficient participant support

A pretrained view. An input bottleneck.

A fixed public EEG encoder runs on nine of 240 candidate segments, with four waveform controls each. The strict quality rules leave no complete eligible participant pairs, so the planned agreement tests cannot be estimated.

All exclusions and segment observations remain visible. The result establishes a reproducible model execution path, without establishing universal tokens or brain meaning.

45 forward passes · 4.5 eligible minutesRead Experiment 009
NOTE 008 / REPEATED SESSIONSFrozen numerical comparison

Different night. The same patterns?

Twelve new ear-EEG recordings add 86.38 qualified hours. Six participants each contribute two nights; the frozen comparison analyzes 48 hours and asks whether session summaries recur across nights.

All four documented ear channels are used. No stage or personal labels enter numerical discovery. Stable person or sensor effects can contribute to recurrence; this does not establish universal tokens.

Four participants remain reserved.Read Experiment 008
NOTES 003–007 / CONTINUOUS CORPUSPublished numerical study

What changes when brainwaves shift?

The full open around-ear inventory contains 55 recordings. Fifty-four qualify, totaling 223.47 sample-hours. We compare the timing and geometry of changes in waveform shape, spectrum and sensor coordination, with phase and nuisance controls.

The methods receive numerical waves without identity or task labels. These measurements test repeatable signal structure; they do not establish a universal brain language. Experiment 005 is a prepared Neurable protocol, with no physical collection yet.

Failures and denominators retained.Read the measured results
NOTE 002 / AROUND-EAR STUDY Analysis in progress

Different views, different vocabularies

Eight around-ear recordings were checked numerically and tokenized with three unsupervised feature families. All 27 seed and vocabulary settings were retained and reproduced.

8fixed-selected recordings
2,769,602,360source bytes
600 sslice per recording, starting at 60 s

VERIFIED OUTCOME

In progress — results will appear after verification.

Measured agreement matrices: spectrum and time-frequency agree more than either agrees with waveform shape. External adjusted Rand scores are 0.289, 0.039 and 0.019; identical partitions would score 1.
Measured Experiment 002 results. Select the graphic to inspect the full-size matrix.
Source selection and slice are fixed.Protocol
NOTE 001 / HISTORICAL SEEDCompleted baseline

A small spectral baseline meets external data

A scalp-trained, single-channel spectral codebook was an early baseline. Its external coverage was low; this does not demonstrate universal patterns or brain meaning.

14recordings
27,266two-second, single-channel windows
18,374tokens
EXTERNAL WINDOW COVERAGE11.36%

393 / 3,458 windows covered. Another 3,011 were out of distribution and 54 were clipped. The denominator counts windows, not people or meaning.

Historical seed · bounded interpretationDownload summary JSON

05 / REPRODUCIBLE METHODS & LIMITS

Make every transformation visible.

The source files remain intact. Derived numeric views, quality decisions, seeds, and run settings belong in the release record with the outcomes they produce.

01

Fix the input

Experiment 002 uses eight selected recordings and the interval from 60 to 660 seconds in each. Preserve original samples; document any re-reference, filter, resampling, and window rule as fixed parameters before reporting results.

02

Learn without labels

Run numeric QC, then learn waveform, spectrum, and time-frequency codebooks with multiple seeds and K = 8, 16, and 32. Discovery sees no hidden task or identity labels; audit metadata remains separate.

03

Test the held-out recordings

Held-out means testing on recordings that did not teach the tokenizer its patterns.

04

Publish the numerical outcome

Report adjusted Rand and adjusted mutual information, which tolerate renamed token IDs; occupied and effective vocabulary; coverage; reconstruction; and shuffled controls. Publish both successes and failures.