Project Description
The Tachyon Resilient Modeling project develops the theories, algorithms, software and datasets needed to make large-scale computational models of data processing pipelines. Our goal is to make these models both predictive and able to explore the resilience of complex designs to improve them against failure during cricital operations.
We use descrete timestep modeling combined with AI/ML surrogate techniques to model different stages and steps in a given workflow. We then perturb these models to explore the operational behavior of the system as a whole. In particular we are interested in realtime triggering and data processing that is used by high energy physics experiments to detect and analyze rare events.
In addition to our modelling work, we are also developing open datasets for model training and evaluation which are based on the real operations of high energy physics experiments, their data analysis workflows, and large scale data center operations. These currated datasets contain not only real world data patterns and operational details, but also include highly detailed information on historic failures. We hope that these datasets will serve as a foundation for future research and benchmarks that can be used by the community to evaluate the resilience of their own models and workflows.
This site collects the project's people, publications, and released research artifacts.
Research Focus Areas
Uncertainty Quantification
Rigorous, scalable propagation of uncertainty through coupled models so predictions come with honest, calibrated error bars.
Fault-Tolerant Methods
Numerical algorithms and workflows that detect, absorb, and recover from hardware, data, and component failures at scale.
Reproducible Workflows
Provenance-tracked, portable modeling pipelines that produce the same results across platforms and over time.
Adaptive Modeling
Models that self-monitor and adjust fidelity as conditions and data availability change during a run.
Open Benchmarks
Curated datasets and reference problems for evaluating resilience and reproducibility across the community.
High-Performance Computing
Implementations tuned for HPC and heterogeneous accelerators with an emphasis on scalability and portability.




