Reference health

3D CNN-Based Speech Emotion Recognition Using K-Means Clustering and Spectrograms

https://doi.org/10.3390/e21050479
CiteStamped reference-health badge
37/37 checkable references clean · checked 2026-07-24

Every reference with a DOI in the deposited reference list resolved to a known work in Crossref or DataCite at the dated check, and none carried a retraction, withdrawal, or removal notice.

7 without a DOI — not checked. A reference deposited without a DOI is never matched by title or guessed at; it stays outside the checked set, and this line discloses that.

The 37 checked references that resolve
resolves10.1007/s10470-017-1006-3
Real-time ensemble based face recognition system for NAO humanoids using local binary pattern
resolves10.1109/ACCESS.2018.2831927
Dominant and Complementary Emotion Recognition From Still Images of Faces
resolves10.1109/CVPR.2015.7298682
FaceNet: A unified embedding for face recognition and clustering
resolves10.1109/T-AFFC.2011.37
Multimodal Emotion Recognition in Response to Videos
resolves10.1007/s12193-009-0025-5
Multimodal emotion recognition in speech-based interaction using facial expression, body gesture and acoustic analysis
resolves10.1007/s10772-017-9396-2
Vocal-based emotion recognition using random forests and decision tree
resolves10.1109/TMM.2014.2360798
Learning Salient Features for Speech Emotion <newline/>Recognition Using Convolutional <newline/>Neural Networks
resolves10.1109/ACCESS.2017.2761539
3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition
resolves10.1007/s00138-018-0960-9
Audiovisual emotion recognition in wild
resolves10.1109/ICASSP.2013.6638346
Deep learning for robust feature generation in audiovisual emotion recognition
resolves10.1109/ICASSP.2011.5947700
Learning a better representation of speech soundwaves using restricted boltzmann machines
resolves10.21437/Interspeech.2015-3
Analysis of CNN-based speech recognition system using raw speech as input
resolves10.1109/PlatCon.2017.7883728
Speech Emotion Recognition from Spectrograms with Deep Convolutional Neural Network
resolves10.1109/TASLP.2014.2339736
Convolutional Neural Networks for Speech Recognition
resolves10.1109/LSP.2010.2100380
Spectrogram Image Feature for Sound Event Classification in Mismatched Conditions
resolves10.1109/TMM.2008.927665
Recognizing Human Emotional State From Audiovisual Signals*
resolves10.1109/ICDEW.2006.145
The eNTERFACE’05 Audio-Visual Emotion Database
resolves10.1007/s11042-016-4041-7
Determining speaker attributes from stress-affected speech in emergency situations with hybrid SVM-DNN architecture
resolves10.1007/s11042-017-5292-7
Deep features-based speech emotion recognition for smart affective services
resolves10.21437/Interspeech.2018-1811
Speech Emotion Recognition Using Spectrogram & Phoneme Embedding
resolves10.1007/s10579-008-9076-6
IEMOCAP: interactive emotional dyadic motion capture database
resolves10.23919/APSIPA.2018.8659587
Attention Based Fully Convolutional Network for Speech Emotion Recognition
resolves10.21437/Interspeech.2005-446
A database of German emotional speech
resolves10.21437/Interspeech.2017-200
Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms
resolves10.1109/TAFFC.2017.2713783
Audio-Visual Emotion Recognition in Video Clips
resolves10.1109/ICSPCS.2010.5709770
Preference for 20-40 ms window duration in speech analysis
resolves10.1109/ICCV.2015.510
Learning Spatiotemporal Features with 3D Convolutional Networks
resolves10.1109/TIT.1982.1056489
Least squares quantization in PCM
resolves10.1109/TIT.2014.2375327
Randomized Dimensionality Reduction for <inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula>-Means Clustering
resolves10.1109/ACII.2017.8273628
Learning spectro-temporal features with 3D CNNs for speech emotion recognition
resolves10.1007/978-3-642-24600-5_64
Audio Visual Emotion Recognition Based on Triple-Stream Dynamic Bayesian Network Models
resolves10.1109/WACV.2017.58
Cyclical Learning Rates for Training Neural Networks
resolves10.1007/s11263-015-0816-y
ImageNet Large Scale Visual Recognition Challenge
resolves10.1121/1.1915715
The Relation of Pitch to Intensity
resolves10.1109/TASSP.1980.1163420
Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
resolves10.1109/ISSPA.2010.5605491
Choice of Mel filter bank in computing MFCC of a resampled speech
The 7 references without a DOI — listed, not checked
no DOI — not checkedSchlüter, J., and Grill, T. (2015, January 26–30). Exploring Data Augmentation for Improved Singing Voice Detection with Neural Networks. Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR 2015), Malaga, Spain.
no DOI — not checkedJackson, P., and Haq, S. (2014). Surrey Audio-Visual Expressed Emotion (SAVEE) Database, University of Surrey.
no DOI — not checkedSimonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv.
no DOI — not checkedKrizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (NIPS 2012), Curran Associates, Inc.
no DOI — not checkedIoffe, S., and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv.
no DOI — not checkedKingma, D.P., and Ba, J.L. (2015, January 7–9). Adam: Amethod for stochastic optimization. Proceedings of the 3rd International Conference for Learning Representations, San Diego, CA, USA.
no DOI — not checkedVidyamurthy, G. (2004). Pairs Trading: Quantitative Methods and Analysis, John Wiley & Sons.
What this badge says. CiteStamped means the CHECKABLE references of this work were clean at the dated check: each resolved to a known work in a public registry, and none carried a retraction notice at that time. It says nothing about the quality, findings, or importance of the work itself, and nothing about references deposited without a DOI.

checked 2026-07-24 — re-checked daily as this page is visited; titles and statuses come from Crossref and DataCite and are not part of the signed record

Embed this badge

Both snippets point at the live badge image and link back to this page. The badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.

<a href="https://citestamp.com/citestamped/10.3390/e21050479"><img src="https://citestamp.com/citestamped/10.3390/e21050479/badge.svg" alt="CiteStamped reference-health badge" width="460" height="64"></a>
[![CiteStamped reference-health badge](https://citestamp.com/citestamped/10.3390/e21050479/badge.svg)](https://citestamp.com/citestamped/10.3390/e21050479)