Resilient Computational Science

Tachyon Resilient Modeling

Tachyon is a Dept. of Energy-funded research project developing methods, tools, and open datasets for building computational models that can model complex domain science workflows. These models are designed to provide multi-scale, multi-fidelity modeling that allows for end-to-end modeling of data analysis pipelines for High Energy Physics.

Our models are designed to allow for reseraches to probe and predict the operational behavior of the data pipelines and to explore deviations from expected behavior to understand fault conditions and to harden our systems against failure modes.

Project Description

The Tachyon Resilient Modeling project develops the theories, algorithms, software and datasets needed to make large-scale computational models of data processing pipelines. Our goal is to make these models both predictive and able to explore the resilience of complex designs to improve them against failure during cricital operations.

We use descrete timestep modeling combined with AI/ML surrogate techniques to model different stages and steps in a given workflow. We then perturb these models to explore the operational behavior of the system as a whole. In particular we are interested in realtime triggering and data processing that is used by high energy physics experiments to detect and analyze rare events.

In addition to our modelling work, we are also developing open datasets for model training and evaluation which are based on the real operations of high energy physics experiments, their data analysis workflows, and large scale data center operations. These currated datasets contain not only real world data patterns and operational details, but also include highly detailed information on historic failures. We hope that these datasets will serve as a foundation for future research and benchmarks that can be used by the community to evaluate the resilience of their own models and workflows.

This site collects the project's people, publications, and released research artifacts.

Research Focus Areas

🌡️

Uncertainty Quantification

Rigorous, scalable propagation of uncertainty through coupled models so predictions come with honest, calibrated error bars.

🛡️

Fault-Tolerant Methods

Numerical algorithms and workflows that detect, absorb, and recover from hardware, data, and component failures at scale.

🔄

Reproducible Workflows

Provenance-tracked, portable modeling pipelines that produce the same results across platforms and over time.

📊

Adaptive Modeling

Models that self-monitor and adjust fidelity as conditions and data availability change during a run.

🧮

Open Benchmarks

Curated datasets and reference problems for evaluating resilience and reproducibility across the community.

High-Performance Computing

Implementations tuned for HPC and heterogeneous accelerators with an emphasis on scalability and portability.

Explore the Project