Reliable AI systems
Infrastructure drift, serving-stack changes, evaluation reliability, and safety-boundary measurement.
DriftBenchSafety framing ablationResearch directions
We study how models, numerical systems, and research infrastructure interact—and build tests that make those interactions inspectable.
Current agenda
These areas are grounded in published and ongoing work by founder Gianluigi Vitale and are open to focused collaboration with other independent researchers.
Infrastructure drift, serving-stack changes, evaluation reliability, and safety-boundary measurement.
DriftBenchSafety framing ablationAccelerator portability, attention implementations, multi-chip TPU bring-up, and long-context execution.
DeepSeek V4-Flash on TPUDSA kernel for TPUBit-level accelerator behavior, deterministic simulation, validation, and reproducibility across numerical paths.
TPU MXU characterizationDSProfilerLong-context representations and scientific questions that require both systems rigor and domain expertise.
Long-context genomicsMethod
State the claim and failure condition before scaling the experiment.
Separate interpolation from true held-out generalization.
Preserve configs, versions, artifacts, and negative results.
Publish only what can be explained and checked.