Research record
Experiments by research track
These are technical research reports, not peer-reviewed publications. Each report identifies its source, measurement conditions and limitations. The evidence section explains what can be downloaded or reproduced.
Understanding time · controlled experiment
The model uses timing information to select the right interval
Does the model use actual timing information to select the correct interval between events, or does it guess from clues in the text?
Experimental criteria met · three training runs
The experiment tests selection of a time interval from given options. It does not establish a ready-to-use language model that can answer any time-related question. An earlier experiment also left reliable matching of repeated intervals to events unresolved.
Read the report →Comparing durations · training experiment
Duration reasoning: from comparison to numerical answers
Does also teaching the model ratios between durations help it identify which duration is longer and estimate its value?
Partial progress · numerical-answer target not met
This is not a completed system for answering duration questions. Further work must align predictions with real durations and limit answers when the model is uncertain.
Read the report →Pyrodit RM · R7 recovery study
Risk detection: training continuity and contextual precision
Can retaining the full model and restoring broader training data improve detection without increasing false alarms or losing existing capabilities?
R7 evaluation complete · R5 retained as the serving baseline
These are English development results, with a separately selected threshold for each evaluation group, not one validated production threshold. They do not establish performance in Turkish, German or other risk categories. The context set had been examined before and is not an untouched independent test. No final context test was run for R7 because no candidate qualified; the reported context figures belong to the separate word-pattern control.
Read the report →Open source · Vector compression
Semafold: smaller vectors, measurable trade-offs
How much space can AI vectors save when the complete compressed representation, including metadata, is counted?
Synthetic benchmark · source report 0.1.0, 31 March 2026 · library reviewed at 0.2.0
Compression is lossy. Real retrieval quality and end-to-end model behaviour need testing on the intended workload. Very small cache blocks can grow relative to 16-bit storage because of metadata overhead. The compression and cache APIs remain preview interfaces; ready-made serving integrations are outside this report.
Read the report →