Reference health

E2EWatch: An End-to-End Anomaly Diagnosis Framework for Production HPC Systems

https://doi.org/10.1007/978-3-030-85665-6_5
CiteStamped reference-health badge
33/33 checkable references clean · checked 2026-07-22

Every reference with a DOI in the deposited reference list resolved to a known work in Crossref or DataCite at the dated check, and none carried a retraction, withdrawal, or removal notice.

4 without a DOI — not checked. A reference deposited without a DOI is never matched by title or guessed at; it stays outside the checked set, and this line discloses that.

The 33 checked references that resolve
resolves10.1109/SC.2014.18
The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications
resolves10.1109/CLUSTER.2015.71
Toward Rapid Understanding of Production HPC Applications and Systems
resolves10.1109/ICCAC.2015.32
Toward Autonomic Cloud: Automatic Anomaly Detection and Resolution
resolves10.1145/2934872.2934884
Taking the Blame Game out of Data Centers Operations with NetPoirot
resolves10.1007/978-3-319-96983-1_7
Taxonomist: Application Detection Through Rich Monitoring Data
resolves10.1145/3203217.3205863
The D.A.V.I.D.E. big-data-powered fine-grain power and performance monitoring support
resolves10.1145/2503210.2503247
There goes the neighborhood
resolves10.1109/IPDPS47924.2020.00096
The Case of Performance Variability on Dragonfly-based Systems
resolves10.1016/j.engappai.2019.07.008
A semisupervised autoencoder-based approach for anomaly detection in high performance computing systems
resolves10.1145/3339186.3339210
Operational Data Analytics
resolves10.1109/NOMS.2016.7502990
Expedite feature extraction for enhanced cloud anomaly detection
resolves10.1109/IPDPS47924.2020.00115
Aarohi: Making Real-Time Node Failure Prediction Feasible
resolves10.1109/IPDPS.2014.27
CALCioM: Mitigating I/O Interference in HPC Systems through Cross-Application Coordination
resolves10.1145/3038912.3052649
Performance Monitoring and Root Cause Analysis for Cloud-hosted Web Applications
resolves10.1109/CLUSTER.2017.23
Data Mining-Based Analysis of HPC Center Operations
resolves10.1109/TPDS.2009.52
Toward Automated Anomaly Identification in Large-Scale Systems
resolves10.2172/918344
Algorithmic support for commodity-based parallel computing systems.
resolves10.1145/3149412.3149421
An empirical survey of performance and energy efficiency variation on Intel processors
resolves10.1080/01621459.1951.10500769
The Kolmogorov-Smirnov Test for Goodness of Fit
resolves10.1016/j.parco.2004.04.001
The ganglia distributed monitoring system: design, implementation, and experience
resolves10.1145/2783258.2788624
Learning a Hierarchical Monitoring System for Detecting and Diagnosing Service Issues
resolves10.1145/3369583.3392674
DCDB Wintermute: Enabling Online and Holistic Operational Data Analytics on HPC Systems
resolves10.1109/CLUSTER49012.2020.00062
HPC System Data Pipeline to Enable Meaningful Insights through Analysis-Driven Visualizations
resolves10.1016/j.procs.2018.08.235
An approach for dynamic detection of inefficient supercomputer applications
resolves10.1109/IISWC.2005.1526010
Understanding the causes of performance variability in HPC workloads
resolves10.1109/TPDS.2018.2870403
Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning
resolves10.1109/TVCG.2018.2865026
A Visual Analytics Framework for the Detection of Anomalous Call Stack Trees in High Performance Computing Applications
resolves10.1007/978-3-319-96983-1_12
Early Termination of Failed HPC Jobs Through Machine and Deep Learning
resolves10.1109/CLOUD.2016.0136
TaskInsight: A Fine-Grained Performance Anomaly Detection and Problem Locating System
resolves10.1109/CLUSTER49012.2020.00026
Quantifying the impact of network congestion on application performance and network metrics
The 4 references without a DOI — listed, not checked
no DOI — not checkedBrandt, J.M., et al.: Enabling advanced operational analysis through multi-subsystem data integration on trinity. Technical report, Sandia National Lab. (SNL-CA), Livermore, CA (United States) (2015)
no DOI — not checkedKe, G., et al.: Lightgbm: a highly efficient gradient boosting decision tree. Adv. Neural. Inf. Process. Syst. 30, 3146–3154 (2017)
no DOI — not checkedPedregosa, F., et al.: Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12, 2825–2830 (2011)
no DOI — not checkedSandia National Laboratories: HPC capacity cluster platforms (2017). https://hpc.sandia.gov/HPC%20Production%20Clusters/index.html
What this badge says. CiteStamped means the CHECKABLE references of this work were clean at the dated check: each resolved to a known work in a public registry, and none carried a retraction notice at that time. It says nothing about the quality, findings, or importance of the work itself, and nothing about references deposited without a DOI.

checked 2026-07-22 — re-checked daily as this page is visited; titles and statuses come from Crossref and DataCite and are not part of the signed record

Embed this badge

Both snippets point at the live badge image and link back to this page. The badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.

<a href="https://citestamp.com/citestamped/10.1007/978-3-030-85665-6_5"><img src="https://citestamp.com/citestamped/10.1007/978-3-030-85665-6_5/badge.svg" alt="CiteStamped reference-health badge" width="460" height="64"></a>
[![CiteStamped reference-health badge](https://citestamp.com/citestamped/10.1007/978-3-030-85665-6_5/badge.svg)](https://citestamp.com/citestamped/10.1007/978-3-030-85665-6_5)