Performance recordings use .sgrec files. They contain separate input stems,
mono sub lanes, quad effect returns, and the actual mastered stereo and quad
outputs. Both masters are recorded before the global listening-volume knob.
Turning that knob down does not quiet the recording. No mastering or spatial
reconstruction is applied during extraction.
Sampler-looper WAV recording is unchanged.
The recording pad blinks red four times per second when an error is latched. The async log records that error once, with its state, accepted/written frame counts, written bytes and queue high-water mark.
Use Python 3.10 or newer. Decoding selected v4 FLAC audio also needs libFLAC
(brew install flac on macOS, or your platform’s libFLAC package). No Python
package is required: the reader uses ctypes. Header inspection, patch queries,
and legacy v1–v3 extraction do not load libFLAC.
python3 scripts/extract_recording.py info session.sgrec --verify
python3 scripts/extract_recording.py extract session.sgrec --master stereo
python3 scripts/extract_recording.py extract session.sgrec --master quad -o quad.wav
python3 scripts/extract_recording.py extract session.sgrec --track 3 -o stem.wav
python3 scripts/extract_recording.py patch session.sgrec --sample 96000 -o patch.json
info lists IDs, types and recording taps. --verify reads every block and the
completion marker. Track selection uses the numeric ID, not its position in the
table. Panned-mono extraction writes just the audio stream as mono; coordinates
remain available through scripts/sgrec.py. Existing outputs are refused unless
--overwrite is supplied.
Exports preserve exact PCM24 integers, sample rate, channel order and frame count. Small outputs use RIFF WAV; large outputs become RF64. The writer reserves the RF64 header space as a JUNK chunk so extraction needs only one pass. Quad WAVs have an unspecified speaker mask and preserve engine order.
On corruption, truncation or missing completion, extraction retains a playable WAV containing only whole validated blocks, reports the frame count and reason, and exits nonzero. An invalid header or unknown selected track creates no WAV. The reader never skips damaged blocks or invents silence for missing blocks. A clean empty recording produces an empty WAV.
scripts/sync_ipad.py detects new files by SMRTGRID magic and extracts their
stereo master. Legacy RIFF/RF64 recordings keep the existing SoX extraction path.
Only a successful, complete extraction permits deletion of the remote original;
a playable partial WAV is insufficient.
patch reconstructs the initial patch plus recorded parameter edits through
--sample, inclusive. Samples are relative to the first recorded audio frame;
96000 is two seconds at 48 kHz. Omitting -o prints JSON to stdout. Existing
output files are refused. The same operation is available as
sgrec.Reader(source).PatchAtSample(sample) on a seekable recording file.
Repeated queries start from the initial patch each time, without an index.
The query validates blocks through the target and does not require a clean tail
after that point. A target past the last audio frame is rejected; sample zero of
a clean empty session returns its initial patch.
The five assignment event types cover StateSaver bytes, gesture faders, scene blend, encoder values, and gesture-encoder activation. These are stored control values before smoothing and modulation, not the resulting DSP output. Replay updates the appropriate patch fields, including nested modulators and gestures created during recording. A gesture’s inherited value is captured separately when activation copies its parent value.
V3 adds PatchLoad and PatchSnapshot events. A load records the supplied patch
and its restoreFaders option as one ordered operation, including removal of
old encoder children. Reset records the resulting live patch as a snapshot.
These operations suppress their internal assignment bursts. Replay preserves
the order of edits before and after each operation, even at the same sample.
This is patch reconstruction, not complete deterministic performance replay. Sample-directory changes and sample assets remain untracked. Ordinary scene copies and value edits use assignment events. A scene switch alone reads stored state and needs no per-value event.
Audio extraction accepts v1 through v4 and ignores patch/event semantics. V4 adds length-prefixed FLAC audio and keeps v3 event semantics. Patch reconstruction requires a v2, v3, or v4 initial patch and recognized event types. V2 retains its historical assignment semantics and cannot reconstruct unrecorded load/reset effects or recover cross-type capture order.
Container integers are little-endian. FLAC payloads follow the native FLAC format, including its big-endian metadata. There is no container alignment or padding.
| Field | Bytes |
|---|---|
ASCII SMRTGRID |
8 |
UTF-8 JSON length (u32) |
4 |
| JSON session header | declared length |
Independent BLK4 records |
variable |
END1 clean-completion marker |
16 |
Required JSON fields:
{
"format_version": 4,
"initial_patch": {"nonagon": {}, "squiggleBoy": {}, "stateSaver": {}, "configGrid": {}, "faders": [], "blend": 0.0},
"recorded_at_utc": "2026-09-14T12:34:56Z",
"git_commit_sha": "0123456789abcdef0123456789abcdef01234567",
"sample_rate": 48000,
"block_frames": 48000,
"tracks": [
{"id": 0, "name": "Voice 1", "type": "panned_mono", "role": "input", "tap": "post_fader"}
]
}
initial_patch contains the complete live patch, including unsaved edits, just
before frame zero. The engine calls ToJSON directly when handling the recording
start, passes the snapshot to the recorder, and immediately begins queuing audio
and subsequent parameter edits. That sample is recording sample zero. The worker
opens the file and serializes the header while capture continues; file opening
and header I/O introduce no gap between the snapshot and recorded data. The date remains UTC session creation time on the worker.
The SHA is the full commit ID
embedded at build time. Runtime recording never invokes Git. No dirty-state
field is recorded.
Track IDs are unique unsigned 32-bit integers. Names, roles and taps are nonempty strings. The table remains fixed for the session. Type determines the streams:
| Type | Ordered streams |
|---|---|
mono |
sample_val |
panned_mono |
sample_val, x, y |
stereo |
left, right |
quad |
q0, q1, q2, q3 |
Quad corners are (0,1), (1,1), (1,0), (0,0). Master roles are uniquely
master_stereo and master_quad, with matching types and tap
post_mastering_pre_master_volume.
Audio conversion preserves the previous WAV writer:
lround(clamp(sample, -1, 1) * 8388607) using a double intermediate. Thus the
writer emits -8388607 through 8388607, while the codec supports the full signed
24-bit range -8388608 through 8388607. There is no dither or normalization.
Coordinates use lround(clamp(position, 0, 1) * 8388607). Divide by 8388607 to
decode a position. Coordinates 0, 0.5 and 1 produce 0, 4194304 and 8388607.
Nonfinite submissions fail capture before accepting that frame.
| Field | Representation |
|---|---|
| Tag | BLK4 (4 bytes) |
| Total record bytes, including CRC | u32 |
| Start frame | u64 |
| Frame count | u32 |
| Included track count | u16 |
| Descriptors, ascending track ID | variable |
| Stream payloads, descriptor order | variable |
| Event group count | u32 |
| Event groups | variable |
| CRC-32/ISO-HDLC | u32 |
The fixed header is 22 bytes. Each descriptor contains track_id:u32, followed
by (encoding:u8, bit_width:u8) for each type-derived stream. No stream count or
legacy payload length is repeated. FLAC payloads carry their own byte length.
Stream payloads start at byte boundaries.
Frame counts range from 1 through the header’s block_frames. Start positions
are contiguous from zero. Only the last data block may be short. Predictors
reset for every stream in every block. Default block size is one second.
A track is omitted when its quantized audio is zero throughout the block.
For panned mono, only sample_val determines omission. Omitted tracks decode to
zero streams, including zero coordinates. An included panned track retains all
coordinates, including those on locally silent frames. Entirely silent blocks
still have records with zero descriptors. With no events they occupy 30 bytes
including CRC; otherwise their event groups follow the empty audio payload.
Before grouping, the writer assigns each event a block-local capture order.
It stable-sorts events by (type, name, sample) and groups by (type, name).
Each group contains:
| Field | Representation |
|---|---|
| Event type | u8; see types below |
| Value width | u8; fixed for the group |
| UTF-8 name length | u16 |
| Entry count | u32 |
| Name | declared UTF-8 bytes, without terminator |
| Entries | type-specific fields below |
Every v3/v4 entry starts with a block-relative sample offset (u32) and capture
order (u32). Replay sorts all entries in the block by (sample, order), so
grouping does not reorder assignments around loads or snapshots. Only fields
relevant to the type follow; there is no struct padding or placeholder data.
| Type | Name | Width | Fields after sample offset and order |
|---|---|---|---|
| 1 StateChange | State name | 1, 2, 4, or 8 | scene:u8, value bytes |
| 2 GestureSet | Empty | 4 | gesture/fader index:u8, value:float32 |
| 3 BlendSet | Empty | 4 | value:float32 |
| 4 EncoderSet | Root parameter name | 4 | scene:u8, track:u8, path length:u8, path bytes, value:float32 |
| 5 EncoderActivate | Root parameter name | 1 | scene:u8, track:u8, path length:u8, path bytes, active:u8 |
| 6 PatchLoad | Empty | 0 | restoreFaders:u8, JSON length:u32, UTF-8 patch JSON |
| 7 PatchSnapshot | Empty | 0 | JSON length:u32, UTF-8 patch JSON |
Floats are finite little-endian IEEE-754 binary32. Scenes are 0..7; tracks and
faders are 0..15. Activation is 0 or 1. Encoder paths have 0..16 hops: each byte
is a modulator index 0..14 or 0x80 | gesture_index (128..143). An empty path
addresses the root encoder; activation requires a final gesture hop. Gestures
are leaves in supported edits; a nested gesture follows zero or more normal
modulator hops. Readers retain support for legacy tagged paths.
The in-memory ParamEvent path uses -1 for unused slots; this sentinel and the
redundant m_gesture/m_isGesture encoder fields are not serialized.
Each name appears once per type per block. GestureSet, BlendSet, PatchLoad, and
PatchSnapshot groups have zero name bytes. Entries retain every submitted value,
including repeated values.
StateChange copies the specified scene’s stored bytes at capture. Replay replaces
bytes at scene * value_width in the named nonagon or stateSaver entry;
wire bytes above 127 become negative byte numbers in patch JSON. It also mirrors
the legacy source-width/selection copies in configGrid.
GestureSet assigns faders[index]; BlendSet assigns top-level blend. Encoder
values are captured in patch units (ToValue), including bipolar values, and
assign squiggleBoy[name]...values.values[scene][track]. Activation assigns the
node’s flat active[scene * 16 + track]. Missing nested nodes start with neutral
zero patch values and inactive gestures; existing nodes and other scenes/tracks
are preserved. The root must already exist in the initial patch. New patches
save and load blend; older patches without it leave the current blend unchanged.
PatchLoad follows the live loader’s partial-load rules. Missing state fields and
encoder roots are preserved. Supplied state byte arrays replace each started
scene value, zero-filling an incomplete final value. A supplied encoder root
loads its values and replaces its child trees, even when modulators or
gestures are omitted. Neutral normal modulators and inactive gestures are
removed as in the engine. configGrid source-width and selection aliases, and
legacy sourceMonitor, overwrite their StateSaver counterparts. A present
configGrid without sourceSelected clears source selections; supplied
selections are normalized to the three-channel limit. Sample directories and
assets are not applied by replay.
restoreFaders is 0 or 1. When 0, both the existing faders and blend survive the
load. When 1, a supplied blend is loaded, and at least 16 supplied faders replace
the current 16 faders; a shorter fader array is ignored. PatchSnapshot replaces
the whole reconstructed patch with the recorded JSON object. Both bulk types
must contain a JSON object; malformed payloads fail patch queries while audio
extraction still skips event contents.
An entry’s timestamp is block_start + sample_offset. Offsets are within the
block’s frame count; an event at the next block boundary belongs to that next
block. Silent and final partial blocks carry events. Events for an audio frame
that was never accepted are excluded.
V1 uses BLK1 and ends immediately after audio payloads and CRC, with a minimum
record size of 26 bytes. V2 uses BLK2, assignment types 1–5, and entries that
start with sample offset alone; it has no capture-order field or bulk patch
types. Readers preserve the stored order of same-sample v2 assignments.
V3 uses BLK3 with the current event representation and legacy audio encodings.
V2 through v4 readers derive audio payload lengths, then skip the remaining
bytes before CRC. They need no event decoder. FLAC stream lengths allow
unselected audio to be skipped after checking its metadata and bounds.
Patch queries decode the trailer and reject an unknown event type in a block needed by the query. They do not decode audio or validate unrelated patch fields.
For N frames:
3*N bytes of little-endian two’s-complement PCM24.3 + ceil((N-1)*width/8). The v4 writer uses this only at width zero for
constant streams, storing one three-byte value. Readers retain all widths.u32 little-endian
byte length precedes a complete native FLAC stream beginning with fLaC.
The length excludes itself. STREAMINFO must declare one channel, 24 bits per
sample, the session sample rate, exactly N samples, and internal block
sizes at most 1024. FLAC metadata must terminate within the payload.
Selected streams are decoded with frame CRC and available MD5 checking,
exact sample-count and consumed-byte checks, and PCM24/coordinate range checks.
Unselected streams require structural metadata and bounds validation, but
do not require the native library or decode frame bodies.The writer uses FLAC level 5 with 1024-sample internal blocks (21.3 ms at 48 kHz). It visits the source once, combines validation and audio activity detection, and feeds the encoder through fixed 1024-sample scratch. FLAC can revisit its bounded internal block. Each scalar stream is independent, including coordinates; there is no stereo or inter-track decorrelation. Encoder allocation and processing run on the existing recording worker.
Whole-constant streams retain the compact three-byte representation. A nonconstant stream of at most 1024 samples uses raw PCM24 when strictly smaller than its length-prefixed FLAC payload, using the already buffered samples. Longer incompressible streams use FLAC’s verbatim subframes. Short final blocks are valid. FLAC predictors and metadata reset at every outer block, preserving independent decoding and the existing one-second recording cadence.
The earlier adaptive-width v4 experiment was never shipped and is superseded by this definition.
CRC covers every preceding byte in the record, using reflected polynomial
0xEDB88320, initial value 0xFFFFFFFF, and final XOR 0xFFFFFFFF. This matches
Python’s zlib.crc32.
The 16-byte completion marker is END1, total_frames:u64, then CRC of those
first 12 bytes. The total must equal the contiguous validated frames. No bytes
may follow it. A zero-frame session contains only the header and completion.
Capture errors drain the accepted prefix where possible and omit completion. Failure reasons remain in runtime status and logs. CRC and completion detect damage; they do not promise power-loss durability. A failed close attempts to remove the completion marker if one was written.
The audio reader validates block framing, audio metadata, IDs, widths, derived
audio sizes, ranges, padding, whole-block CRC, frame continuity and completion.
It parses the header JSON but does not validate or interpret initial_patch or
the event trailer. Patch queries check event framing and the fields they need to update, without a whole-patch schema validator.
Allocation limits are 1 MiB header JSON, 128 tracks, 64 MiB encoded record and 64 MiB
decoded int32 block. Discard each Reader.Blocks() result before requesting the
next one to retain the streaming memory bound.
Prepare allocates storage and starts the persistent worker outside audio.
One SPSC ring holds 32 pages of 1024 frame-major float frames. The worker retains
each slot until consumed, transposes and quantizes into a one-second block, then
compresses and writes directly. Capture only checks finiteness, stages/copies
floats and publishes pages. A second preallocated SPSC queue holds 1024 fixed
parameter events. Events are published before their audio pages; the worker drains
them before finalizing each block, then sorts and groups them. Queue exhaustion
is a recording error, including while the file is opening. Start and stop do no
file work or joining. Even an immediate stop retains the accepted frames; an
empty session writes its header and zero-frame completion marker.
Source-monitor toggles use the ordinary sourceMonitor_i StateChange events and
replay into the global stateSaver section. New patches keep no separate monitor
array in configGrid; legacy patches with that array remain loadable.
A separate 8 MiB JSON arena holds the initial patch snapshot through worker closure, independent of ordinary patch saving. Snapshot construction and event capture allocate no heap memory on audio; JSON text serialization runs on the worker. The engine drains recording before destroying its registered states and root encoder names.
Each queued load retains its load/save PatchArena. The owner cannot reset or
grow that arena until every reader releases it. The worker serializes the patch
as it drains events, releases the arena immediately, and retains its own bytes
until block writing. Busy message-thread loads are deferred before parsing and
retried by the timer; a newer pending request replaces an older pending request.
Save requests and saved-pad reloads defer to a later control frame if necessary.
Error, overflow, shutdown, and discarded-tail paths release retained references.
Whole-patch reset uses another preallocated 8 MiB arena for its resulting snapshot. Reset remains pending while that arena is still referenced. Capture publishes one load or reset operation instead of one event for every affected state and encoder value. Worker-owned payload bytes count toward the 64 MiB pending-event bound; exhaustion fails recording through the explicit error path.
Stop publishes the partial page and a separate stop flag. Ring exhaustion fails capture without waiting or overwriting. New starts are rejected until worker closure is acknowledged. Shutdown runs after audio callbacks are quiescent and drains/joins before storage destruction.
The current mixer has 31 tracks / 78 streams: 17 panned inputs, 9 mono lanes, 3 quad returns and the two masters. Its ring holds about 0.68 seconds; storage is about 35 MiB for audio, plus two 8 MiB arenas for initial/reset snapshots and event storage, before allocator/logging overhead. A longer disk stall can produce an incomplete recording and visible error. Source-width switches keep track IDs and types. Noise mode zeroes skipped stems and still records masters.
Build and run the standalone recorder measurement after configuring CMake tests:
c++ -std=c++17 -O2 -Iprivate/src -I/tmp/smartgrid-recording-build/generated \
-DSMARTGRID_RECORDING_SYSTEM_FLAC=1 $(pkg-config --cflags flac) \
private/test/tools/benchmark_recording.cpp private/src/RecordingFlacCodec.cpp \
$(pkg-config --libs flac) -o /tmp/benchmark_recording -pthread
/tmp/benchmark_recording /tmp/recording-bench audio 30
Modes are silence, audio (sine audio with moving pans), and noise (independent
random audio and coordinates). It writes at real-time rate and reports size,
maximum block encode/write time, queue high-water, capture cost, capture-thread
allocation count, peak RSS (bytes on macOS), and errors. This isolates recorder
cost; it is not a measurement of the whole synth callback or iPad storage.
See the change’s verification.md for measured results and remaining platform
checks. Golden records and their independent construction are in
private/test/fixtures/streaming-recording/.
Standalone C++ builds require pkg-config and the libFLAC development package
(brew install pkg-config flac on macOS). CMake links the system library for
tests. Apple app builds reuse FLAC from the pinned JUCE dependency and do not
link an external/system FLAC installation. Encoder integration lives in
RecordingFlacCodec.cpp; the rest of the core has no FLAC or JUCE headers.
Real-recording comparisons and the band-limited-square investigation are recorded
in docs/research/2026-09-23-sgrec-real-recording-codecs.md. Host timing does not
measure the whole synthesizer or iPad storage performance.