Reference health

The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation

https://doi.org/10.1186/s12864-019-6413-7
CiteStamped reference-health badge
75/75 checkable references clean · checked 2026-07-24

Every reference with a DOI in the deposited reference list resolved to a known work in Crossref or DataCite at the dated check, and none carried a retraction, withdrawal, or removal notice.

34 without a DOI — not checked. A reference deposited without a DOI is never matched by title or guessed at; it stays outside the checked set, and this line discloses that.

The 75 checked references that resolve
resolves10.1371/journal.pone.0208737
Computational prediction of diagnosis and feature selection on mesothelioma patient health records
resolves10.7717/peerj-cs.154
Supervised deep learning embeddings for the prediction of cervical cancer diagnosis
resolves10.1371/journal.pone.0208924
Distillation of the clinical algorithm improves prognosis by multi-task deep learning in high-risk Neuroblastoma
resolves10.1186/s12859-018-2033-5
Phylogenetic convolutional neural networks in metagenomics
resolves10.1038/nature14539
Deep learning
resolves10.4249/scholarpedia.1883
K-nearest neighbor
resolves10.1109/5254.708428
Support vector machines
resolves10.1023/A:1010933404324
Random Forests
resolves10.2741/2712
Classification algorithms for phenotype prediction in genomics and proteomics
resolves10.1093/bioinformatics/btp331
Predictor correlation impacts machine learning algorithms: implications for genomic studies
resolves10.1101/168419
Virtual ChIP-seq: predicting transcription factor binding by learning from the transcriptome
resolves10.1038/ng.3539
Enhancer–promoter interactions are encoded by complex genomic signatures on looping chromatin
resolves10.1093/bioinformatics/btm026
De novo SVM classification of precursor microRNAs from genomic pseudo hairpins using global and intrinsic folding measures
resolves10.1016/j.ipm.2009.03.002
A systematic analysis of performance measures for classification tasks
resolves10.1016/j.patrec.2008.08.010
An experimental comparison of performance measures for classification
resolves10.1109/icpr.2010.156
Theoretical Analysis of a Performance Measure for Imbalanced Data
resolves10.1017/CBO9780511921803
Evaluating Learning Algorithms
resolves10.1186/1471-2164-13-S4-S2
How to evaluate performance of prediction methods? Measures and their interpretation in variation effect analysis
resolves10.17148/IJARCCE.2016.5890
Comparison of the Performance Evaluations in Classification
resolves10.1145/2907070
A Survey of Predictive Modeling on Imbalanced Domains
resolves10.1016/j.chemolab.2017.12.004
Multivariate comparison of classification performance measures
resolves10.1016/j.aci.2018.08.003
Classification assessment methods
resolves10.1016/j.patcog.2019.02.023
The impact of class imbalance in classification performance metrics based on the binary confusion matrix
resolves10.1109/icdm.2011.21
An Analysis of Performance Measures for Binary Classifiers
resolves10.1109/TCBB.2007.1006
Accurate Cancer Classification Using Expressions of Very Few Genes
resolves10.1016/0005-2795(75)90109-9
Comparison of the predicted and observed secondary structure of T4 phage lysozyme
resolves10.1093/bioinformatics/16.5.412
Assessing the accuracy of prediction algorithms for classification: an overview
resolves10.1016/j.compbiolchem.2004.09.006
Comparing two K-category assignments by a K-category correlation coefficient
resolves10.1038/nbt.1665
The MicroArray Quality Control (MAQC)-II study of common practices for the development and validation of microarray-based predictive models
resolves10.1038/nbt.2957
A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium
resolves10.14257/ijhit.2015.8.1.14
Research on the Matthews Correlation Coefficients Metrics of Personalized Recommendation Algorithm Evaluation
resolves10.18632/oncotarget.20923
Precision and recall oncology: combining multiple gene mutations for improved identification of drug-sensitive tumours
resolves10.1002/minf.201700127
Classifiers and their Metrics Quantified
resolves10.1371/journal.pone.0177678
Optimal classifier for imbalanced data using Matthews Correlation Coefficient metric
resolves10.1002/(SICI)1097-4571(199401)45:1<12::AID-ASI2>3.0.CO;2-L
The relationship between Recall and Precision
resolves10.1371/journal.pone.0118432
The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets
resolves10.2307/1932409
Measures of the Amount of Ecologic Association Between Species
resolves10.1108/eb026584
FOUNDATION OF EVALUATION
resolves10.1109/42.363096
Morphometric analysis of white matter lesions in MR images: method and validation
resolves10.1016/0306-4573(92)90005-K
The pragmatics of information retrieval experimentation, revisited
resolves10.3115/112405.112471
Evaluating text categorization
resolves10.1109/TKDE.2010.164
Random k-Labelsets for Multilabel Classification
resolves10.1016/j.patcog.2016.08.008
Designing multi-label classifiers that maximize F measures: State of the art
resolves10.1197/jamia.M1733
Agreement, the F-Measure, and Reliability in Information Retrieval
resolves10.1007/s11222-017-9746-6
A note on using the F-measure for evaluating record linkage algorithms
resolves10.1371/journal.pcbi.1006625
Local epigenomic state cannot discriminate interacting and non-interacting enhancer–promoter pairs with high accuracy
resolves10.1177/001316446002000104
A Coefficient of Agreement for Nominal Scales
resolves10.2307/2529310
The Measurement of Observer Agreement for Categorical Data
resolves10.11613/BM.2012.031
Interrater reliability: the kappa statistic
resolves10.1002/pst.1659
The disagreeable behaviour of the kappa statistic
resolves10.1016/j.eswa.2006.10.022
Comparison of classification accuracy using Cohen’s Weighted Kappa
resolves10.1016/S0031-3203(02)00257-1
Strategies for learning in class imbalance problems
resolves10.1016/j.eswa.2009.11.040
A novel measure for evaluating classifiers
resolves10.1371/journal.pone.0210264
Enhancing Confusion Entropy (CEN) for binary and multiclass classification
resolves10.1371/journal.pone.0041882
A Comparison of MCC and CEN Error Measures in Multi-Class Prediction
resolves10.1109/icpr.2010.764
The Balanced Accuracy and Its Posterior Distribution
resolves10.1148/radiology.143.1.7063747
The meaning and use of the area under a receiver operating characteristic (ROC) curve.
resolves10.1016/S0031-3203(96)00142-2
The use of the area under the ROC curve in the evaluation of machine learning algorithms
resolves10.1109/TKDE.2005.50
Using AUC and accuracy in evaluating learning algorithms
resolves10.1016/j.patrec.2005.10.010
An introduction to ROC analysis
resolves10.1002/sim.3859
Evaluating diagnostic tests: The area under the ROC curve and the balance of errors
resolves10.1111/j.1466-8238.2007.00358.x
AUC: a misleading measure of the performance of predictive distribution models
resolves10.1093/bioinformatics/btq037
Small-sample precision of ROC-related estimates
resolves10.1007/s10994-009-5119-5
Measuring classifier performance: a coherent alternative to the area under the ROC curve
resolves10.1371/journal.pone.0092209
Area under Precision-Recall Curves for Weighted and Unweighted Data
resolves10.1016/j.jclinepi.2015.02.010
The precision–recall curve overcame the optimism of the receiver operating characteristic curve in rare diseases
resolves10.1186/1471-2105-11-523
Class prediction for high-dimensional class-imbalanced data
resolves10.1136/bmj.e4483
Pearson's correlation coefficient
resolves10.2478/v10117-011-0021-1
Comparison of Values of Pearson's and Spearman's Correlation Coefficients on the Same Sets of Data
resolves10.1007/978-3-319-24462-4_2
Extended Spearman and Kendall Coefficients for Gene Annotation List Correlation
resolves10.1073/pnas.96.12.6745
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays
resolves10.1093/bib/bbl016
Partial least squares: a versatile tool for the analysis of high-dimensional genomic data
resolves10.1016/S0167-9473(01)00065-2
Stochastic gradient boosting
resolves10.1007/3-540-49257-7_15
When Is “Nearest Neighbor” Meaningful?
The 34 references without a DOI — listed, not checked
no DOI — not checkedDemšar J. Statistical comparisons of classifiers over multiple data sets,. J Mach Learn Res. 2006; 7:1–30.
no DOI — not checkedGarcía S, Herrera F. An extension on ”Statistical comparisons of classifiers over multiple data sets” for all pairwise comparisons. J Mach Learn Res. 2008; 9:2677–94.
no DOI — not checkedChoi S-S, Cha S-H. A survey of binary similarity and distance measures. J Syst Cybernet Informa. 2010; 8(1):43–8.
no DOI — not checkedPowers DMW. Evaluation: from precision, recall and F-measure to ROC, informedness, markedness & correlation. J Mach Learn Technol. 2011; 2(1):37–63.
no DOI — not checkedAnagnostopoulos C, Hand DJ, Adams NM. Measuring Classification Performance: the hmeasure Package. Technical report, CRAN. 2019:1–17.
no DOI — not checkedSokolova M, Japkowicz N, Szpakowicz S. Beyond accuracy, F-score and ROC: a family of discriminant measures for performance evaluation. In: Proceedings of Advances in Artificial Intelligence (AI 2006), Lecture Notes in Computer Science, vol. 4304. Heidelberg: Springer: 2006. p. 1015–21.
no DOI — not checkedGu Q, Zhu L, Cai Z. Evaluation measures of the classification performance of imbalanced data sets. In: Proceedings of ISICA 2009 – the 4th International Symposium on Computational Intelligence and Intelligent Systems, Communications in Computer and Information Science, vol. 51. Heidelberg: Springer: 2009. p. 461–71.
no DOI — not checkedBekkar M, Djemaa HK, Alitouche TA. Evaluation measures for models assessment over imbalanced data sets. J Informa Eng Appl. 2013; 3(10):27–38.
no DOI — not checkedAkosa JS. Predictive accuracy: a misleading performance measure for highly imbalanced data. In: Proceedings of the SAS Global Forum 2017 Conference. Cary, North Carolina: SAS Institute Inc.: 2017. p. 942–2017.
no DOI — not checkedGuilford JP. Psychometric Methods. New York City: McGraw-Hill; 1954.
no DOI — not checkedCramér H. Mathematical Methods of Statistics. Princeton: Princeton University Press; 1946.
no DOI — not checkedSørensen T. A method of establishing groups of equal amplitude in plant sociology based on similarity of species and its application to analyses of the vegetation on Danish commons. K Dan Vidensk Sels. 1948; 5(4):1–34.
no DOI — not checkedvan Rijsbergen CJ, Joost C. Information Retrieval. New York City: Butterworths; 1979.
no DOI — not checkedChinchor N. MUC-4 evaluation metrics. In: Proceedings of MUC-4 – the 4th Conference on Message Understanding. McLean: Association for Computational Linguistics: 1992. p. 22–9.
no DOI — not checkedTague-Sutcliffe J. The pragmatics of information retrieval experimentation. In: Information Retrieval Experiment, Chap. 5. Amsterdam: Butterworths: 1981.
no DOI — not checkedLewis DD, Yang Y, Rose TG, Li F. RCV1: a new benchmark collection for text categorization research. J Mach Learn Res. 2004; 5:361–97.
no DOI — not checkedLipton ZC, Elkan C, Naryanaswamy B. Optimal thresholding of classifiers to maximize F1 measure. In: Proceedings of ECML PKDD 2014 – the 2014 Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Lecture Notes in Computer Science, vol. 8725. Heidelberg: Springer: 2014. p. 225–39.
no DOI — not checkedSasaki Y. The truth of the F-measure. Teach Tutor Mater. 2007; 1(5):1–5.
no DOI — not checkedPowers DMW. What the F-measure doesn’t measure...: features, flaws, fallacies and fixes. arXiv:1503.06410. 2015.
no DOI — not checkedVan Asch V. Macro-and micro-averaged evaluation measures. Technical report. 2013:1–27.
no DOI — not checkedFlach PA, Kull M. Precision-Recall-Gain curves: PR analysis done right. In: Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 2015). Cambridge: MIT Press: 2015. p. 838–46.
no DOI — not checkedYedidia A. Against the F-score. 2016. Blogpost: https://adamyedidia.files.wordpress.com/2014/11/f_score.pdf. Accessed 10 Dec 2019.
no DOI — not checkedPowers DMW. The problem with Kappa. In: Proceedings of EACL 2012 – the 13th Conference of the European Chapter of the Association for Computational Linguistics. Avignon: ACL: 2012. p. 345–55.
no DOI — not checkedDelgado R, Tibau X-A. Why Cohen’s Kappa should be avoided as performance measure in classification. PloS ONE. 2019; 14(9):0222916.
no DOI — not checkedSebastiani F. An axiomatically derived measure for the evaluation of classification algorithms. In: Proceedings of ICTIR 2015 – the ACM SIGIR 2015 International Conference on the Theory of Information Retrieval. New York City: ACM: 2015. p. 11–20.
no DOI — not checkedEspíndola R, Ebecken N. On extending F-measure and G-mean metrics to multi-class problems. WIT Trans Inf Commun Technol. 2005; 35:25–34.
no DOI — not checkedDubey A, Tarar S. Evaluation of approximate rank-order clustering using Matthews correlation coefficient. Int J Eng Adv Technol. 2018; 8(2):106–13.
no DOI — not checkedFlach PA. The geometry of ROC space: understanding machine learning metrics through ROC isometrics. In: Proceedings of ICML 2003 – the 20th International Conference on Machine Learning. Palo Alto: AAAI Press: 2003. p. 194–201.
no DOI — not checkedSuresh Babu N. Various performance measures in binary classification – An overview of ROC study. Int J Innov Sci Eng Technol. 2015; 2(9):596–605.
no DOI — not checkedFerri C, Hernández-Orallo J, Flach PA. A coherent interpretation of AUC as a measure of aggregated classification performance. In: Proceedings of ICML 2011 – the 28th International Conference on Machine Learning. Norristown: Omnipress: 2011. p. 657–64.
no DOI — not checkedChicco D. Ten quick tips for machine learning in computational biology. BioData Min. 2017; 10(35):1–17.
no DOI — not checkedBoulesteix A-L, Durif G, Lambert-Lacroix S, Peyre J, Strimmer K. Package ‘plsgenomics’. 2018. https://cran.r-project.org/web/packages/plsgenomics/index.html. Accessed 10 Dec 2019.
no DOI — not checkedAlon U, Barkai N, Notterman DA, Gish K, Ybarra S, Mack D, Levine AJ. Data pertaining to the article ‘Broad patterns of gene expression revealed by clustering of tumor and normal colon tissues probed by oligonucleotide arrays’. 2000. http://genomics-pubs.princeton.edu/oncology/affydata/index.html. Accessed 10 Dec 2019.
no DOI — not checkedTimofeev R. Classification and regression trees (CART) theory and applications. Berlin: Humboldt University; 2004.
What this badge says. CiteStamped means the CHECKABLE references of this work were clean at the dated check: each resolved to a known work in a public registry, and none carried a retraction notice at that time. It says nothing about the quality, findings, or importance of the work itself, and nothing about references deposited without a DOI.

checked 2026-07-24 — re-checked daily as this page is visited; titles and statuses come from Crossref and DataCite and are not part of the signed record

Embed this badge

Both snippets point at the live badge image and link back to this page. The badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.

<a href="https://citestamp.com/citestamped/10.1186/s12864-019-6413-7"><img src="https://citestamp.com/citestamped/10.1186/s12864-019-6413-7/badge.svg" alt="CiteStamped reference-health badge" width="460" height="64"></a>
[![CiteStamped reference-health badge](https://citestamp.com/citestamped/10.1186/s12864-019-6413-7/badge.svg)](https://citestamp.com/citestamped/10.1186/s12864-019-6413-7)