The data portal sometimes lists a wide variety of subtypes of datasets pertaining to many machine learning applications.
Datasets from physical systems.
Datasets from biological systems.
This section includes datasets that deals with structured data.
This section includes datasets that ...
As datasets come in myriad formats and can sometimes be difficult to use, there has been considerable work put into curating and standardizing the format of datasets to make them easier to use for machine learning research.
Wissner-Gross, A. "Datasets Over Algorithms". Edge.com. Retrieved 8 January 2016. https://edge.org/response-detail/26587
Weiss, G. M.; Provost, F. (October 2003). "Learning When Training Data are Costly: The Effect of Class Distribution on Tree Induction". Journal of Artificial Intelligence Research. 19: 315–354. doi:10.1613/jair.1199. /wiki/Doi_(identifier)
Abney, Steven (2007). Semisupervised Learning for Computational Linguistics. CRC Press. ISBN 978-1-4200-1080-0.[page needed] 978-1-4200-1080-0
Žliobaitė, Indrė; Bifet, Albert; Pfahringer, Bernhard; Holmes, Geoff (2011). "Active Learning with Evolving Streaming Data". Machine Learning and Knowledge Discovery in Databases. Lecture Notes in Computer Science. Vol. 6913. pp. 597–612. doi:10.1007/978-3-642-23808-6_39. ISBN 978-3-642-23807-9. 978-3-642-23807-9
James Bennett; Stan Lanning (12 August 2007). "The Netflix Prize" (PDF). Proceedings of KDD Cup and Workshop 2007. Archived from the original (PDF) on 27 September 2007. Retrieved 25 August 2007. https://web.archive.org/web/20070927051207/http://www.netflixprize.com/assets/NetflixPrizeKDD_to_appear.pdf
McAuley, Julian; Targett, Christopher; Shi, Qinfeng; Anton van den Hengel (2015). "Image-based Recommendations on Styles and Substitutes". arXiv:1506.04757 [cs.CV]. /wiki/ArXiv_(identifier)
"Amazon review data". nijianmo.github.io. Retrieved 8 October 2021. https://nijianmo.github.io/amazon/index.html
Ganesan, Kavita; Zhai, Chengxiang (2012). "Opinion-based entity ranking". Information Retrieval. 15 (2): 116–150. doi:10.1007/s10791-011-9174-8. hdl:2142/15252. S2CID 16258727. /wiki/Doi_(identifier)
Lv, Yuanhua; Lymberopoulos, Dimitrios; Wu, Qiang (2012). "An exploration of ranking heuristics in mobile local search". Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval. pp. 295–304. doi:10.1145/2348283.2348325. ISBN 978-1-4503-1472-5. 978-1-4503-1472-5
Harper, F. Maxwell; Konstan, Joseph A. (2015). "The MovieLens Datasets: History and Context". ACM Transactions on Interactive Intelligent Systems. 5 (4): 19. doi:10.1145/2827872. S2CID 16619709. /wiki/Doi_(identifier)
Koenigstein, Noam; Dror, Gideon; Koren, Yehuda (2011). "Yahoo! Music recommendations: Modeling music ratings with temporal dynamics and item taxonomy". Proceedings of the fifth ACM conference on Recommender systems. pp. 165–172. doi:10.1145/2043932.2043964. ISBN 978-1-4503-0683-6. 978-1-4503-0683-6
McFee, Brian; Bertin-Mahieux, Thierry; Ellis, Daniel P.W.; Lanckriet, Gert R.G. (2012). "The million song dataset challenge". Proceedings of the 21st International Conference on World Wide Web. pp. 909–916. doi:10.1145/2187980.2188222. ISBN 978-1-4503-1230-1. 978-1-4503-1230-1
Bohanec, Marko, and Vladislav Rajkovic. "Knowledge acquisition and explanation for multi-attribute decision making." 8th Intl Workshop on Expert Systems and their Applications. 1988. https://www.researchgate.net/profile/Marko_Bohanec/publication/246614940_KNOWLEDGE_ACQUISITION_AND_EXPLANATION_FOR_MULTI-ATTRIBUTE_DECISION_MAKING/links/02e7e532152f452d87000000.pdf
Tan, Peter J., and David L. Dowe. "MML inference of decision graphs with multi-way joins." Australian Joint Conference on Artificial Intelligence. 2002. http://www.csse.monash.edu.au/~dld/Publications/2002/Tan+Dowe2002_MMLDecisionGraphs.ps
"Quantifying comedy on YouTube: why the number of o's in your LOL matter". Metatext NLP Database. Retrieved 26 October 2020. https://metatext.io/datasets
Kim, Byung Joo (2012). "A Classifier for Big Data". Convergence and Hybrid Information Technology. Communications in Computer and Information Science. Vol. 310. pp. 505–512. doi:10.1007/978-3-642-32692-9_63. ISBN 978-3-642-32691-2. 978-3-642-32691-2
Pérezgonzález, Jose D.; Gilbey, Andrew (2011). "Predicting Skytrax airport rankings from customer reviews". Journal of Airport Management. 5 (4): 335–339. doi:10.69554/RFZC4321. https://www.ingentaconnect.com/content/hsp/cam/2011/00000005/00000004/art00007
Loh, Wei-Yin, and Yu-Shan Shih. "Split selection methods for classification trees." Statistica sinica(1997): 815–840. http://www3.stat.sinica.edu.tw/statistica/oldpdf/A7n41.pdf
Lim, Tjen-Sien; Loh, Wei-Yin; Shih, Yu-Shan (2000). "A comparison of prediction accuracy, complexity, and training time of thirty-three old and new classification algorithms". Machine Learning. 40 (3): 203–228. doi:10.1023/a:1007608224229. S2CID 17030953. /wiki/Doi_(identifier)
Nguyen, Kiet Van; Nguyen, Vu Duc; Nguyen, Phu X. V.; Truong, Tham T. H.; Nguyen, Ngan Luu-Thuy (2018). "UIT-VSFC: Vietnamese Students' Feedback Corpus for Sentiment Analysis". 2018 10th International Conference on Knowledge and Systems Engineering (KSE). pp. 19–24. doi:10.1109/KSE.2018.8573337. ISBN 978-1-5386-6113-0. 978-1-5386-6113-0
Ho, Vong Anh; Nguyen, Duong Huynh-Cong; Nguyen, Danh Hoang; Pham, Linh Thi-Van; Nguyen, Duc-Vu; Nguyen, Kiet Van; Nguyen, Ngan Luu-Thuy (2020). "Emotion Recognition for Vietnamese Social Media Text". Computational Linguistics. Communications in Computer and Information Science. Vol. 1215. pp. 319–333. arXiv:1911.09339. doi:10.1007/978-981-15-6168-9_27. ISBN 978-981-15-6167-2. S2CID 208202333. 978-981-15-6167-2
Nhung Thi-Hong Nguyen, Phuong Ha-Dieu Phan, Luan Thanh Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen (24 April 2021). "Vietnamese Open-domain Complaint Detection in E-Commerce Websites". arXiv:2104.11969 [cs.CL].{{cite arXiv}}: CS1 maint: multiple names: authors list (link) /wiki/ArXiv_(identifier)
Phu Gia Hoang, Canh Duc Luu, Khanh Quoc Tran, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen (26 January 2023). "ViHOS: Hate Speech Spans Detection for Vietnamese". arXiv:2301.10186 [cs.CL].{{cite arXiv}}: CS1 maint: multiple names: authors list (link) /wiki/ArXiv_(identifier)
Dermouche, Mohamed; Velcin, Julien; Khouas, Leila; Loudcher, Sabine (2014). "A Joint Model for Topic-Sentiment Evolution over Time". 2014 IEEE International Conference on Data Mining. IEEE. pp. 773–778. doi:10.1109/icdm.2014.82. ISBN 978-1-4799-4302-9. 978-1-4799-4302-9
Rose, Tony; Stevenson, Mark; Whitehead, Miles (2002). "The Reuters Corpus Volume 1-from Yesterday's News to Tomorrow's Language Resources". LREC. 2. S2CID 9239414. /wiki/S2CID_(identifier)
Amini, Massih R.; Usunier, Nicolas; Goutte, Cyril (2009). "Learning from Multiple Partially Observed Views – an Application to Multilingual Text Categorization". Advances in Neural Information Processing Systems. 22: 28–36. http://papers.nips.cc/paper/3690-learning-from-multiple-partially-observed-views-an-application-to-multilingual-text-categorization
Liu, Ming; et al. (2015). "VRCA: a clustering algorithm for massive amount of texts". Proceedings of the 24th International Conference on Artificial Intelligence. AAAI Press. Archived from the original on 5 November 2021. Retrieved 6 August 2019. https://web.archive.org/web/20211105004605/https://www.aaai.org/ocs/index.php/IJCAI/IJCAI15/paper/download/10903/10990
Al-Harbi, S; Almuhareb, A; Al-Thubaity, A; Khorsheed, M. S.; Al-Rajeh, A (2008). "Automatic Arabic Text Classification". Proceedings of the 9th International Conference on the Statistical Analysis of Textual Data, Lyon, France.
"Relationship and Entity Extraction Evaluation Dataset: Dstl/re3d". GitHub. 17 December 2018. https://github.com/dstl/re3d
"The Examiner – SpamClickBait Catalogue". https://www.kaggle.com/therohk/examine-the-examiner
"A Million News Headlines". https://www.kaggle.com/therohk/million-headlines
"One Week of Global News Feeds". https://www.kaggle.com/therohk/global-news-week
Kulkarni, Rohit (2018), Reuters News-Wire Archive, Harvard Dataverse, doi:10.7910/DVN/XDB74W /wiki/Doi_(identifier)
"IrishTimes – the Waxy-Wany News". https://www.kaggle.com/therohk/ireland-historical-news
"News Headlines Dataset For Sarcasm Detection". kaggle.com. Retrieved 27 April 2019. https://kaggle.com/rmisra/news-headlines-dataset-for-sarcasm-detection
Klimt, Bryan, and Yiming Yang. "Introducing the Enron Corpus." CEAS. 2004. https://bklimt.com/papers/2004_klimt_ceas.pdf
Kossinets, Gueorgi; Kleinberg, Jon; Watts, Duncan (2008). "The Structure of Information Pathways in a Social Communication Network". arXiv:0806.3201 [physics.soc-ph]. /wiki/ArXiv_(identifier)
Androutsopoulos, Ion; Koutsias, John; Chandrinos, Konstantinos V.; Paliouras, George; Spyropoulos, Constantine D. (2000). "An evaluation of Naive Bayesian anti-spam filtering". In Potamias, G.; Moustakis, V.; van Someren, M. (eds.). Proceedings of the Workshop on Machine Learning in the New Information Age. 11th European Conference on Machine Learning, Barcelona, Spain. Vol. 11. pp. 9–17. arXiv:cs/0006013. Bibcode:2000cs........6013A. /wiki/ArXiv_(identifier)
Bratko, Andrej; et al. (2006). "Spam filtering using statistical data compression models" (PDF). The Journal of Machine Learning Research. 7: 2673–2698. http://www.jmlr.org/papers/volume7/bratko06a/bratko06a.pdf
Almeida, Tiago A., José María G. Hidalgo, and Akebo Yamakami. "Contributions to the study of SMS spam filtering: new collection and results."Proceedings of the 11th ACM symposium on Document engineering. ACM, 2011. http://www.dt.fee.unicamp.br/~tiago/smsspamcollection/doceng11.pdf
Delany; Jane, Sarah; Buckley, Mark; Greene, Derek (2012). "SMS spam filtering: methods and data". Expert Systems with Applications. 39 (10): 9899–9908. doi:10.1016/j.eswa.2012.02.053. S2CID 15546924. https://arrow.dit.ie/cgi/viewcontent.cgi?article=1022&context=scschcomart
Joachims, Thorsten. A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization. No. CMU-CS-96-118. Carnegie-mellon univ pittsburgh pa dept of computer science, 1996. https://apps.dtic.mil/dtic/tr/fulltext/u2/a307731.pdf
Dimitrakakis, Christos, and Samy Bengio. Online Policy Adaptation for Ensemble Algorithms. No. EPFL-REPORT-82788. IDIAP, 2002. https://infoscience.epfl.ch/record/82788/files/rr02-28.pdf
Dooms, S. et al. "Movietweetings: a movie rating dataset collected from twitter, 2013. Available from https://github.com/sidooms/MovieTweetings." https://github.com/sidooms/MovieTweetings
RoyChowdhury, Aruni; Lin, Tsung-Yu; Maji, Subhransu; Learned-Miller, Erik (2017). "Twitter100k: A Real-world Dataset for Weakly Supervised Cross-Media Retrieval". arXiv:1703.06618 [cs.CV]. /wiki/ArXiv_(identifier)
"huyt16/Twitter100k". GitHub. Retrieved 26 March 2018. https://github.com/huyt16/Twitter100k
Go, Alec; Bhayani, Richa; Huang, Lei (2009). "Twitter sentiment classification using distant supervision". CS224N Project Report, Stanford. 1: 12.
Chikersal, Prerna, Soujanya Poria, and Erik Cambria. "SeNTU: sentiment analysis of tweets by combining a rule-based classifier with supervised learning." Proceedings of the International Workshop on Semantic Evaluation, SemEval. 2015. https://www.aclweb.org/anthology/S15-2108
Zafarani, Reza, and Huan Liu. "Social computing data repository at ASU." School of Computing, Informatics and Decision Systems Engineering, Arizona State University (2009). /wiki/Huan_Liu
Data Science Course by DataTrained Education "IBM Certified Data Science Course." IBM Certified Online Data Science Course https://web.archive.org/web/20200928084958/https://www.datatrained.com/data-science-program
McAuley, Julian J.; Leskovec, Jure. "Learning to Discover Social Circles in Ego Networks". NIPS. 2012: 2012.
Šubelj, Lovro; Fiala, Dalibor; Bajec, Marko (2014). "Network-based statistical comparison of citation topology of bibliographic databases". Scientific Reports. 4 (6496): 6496. arXiv:1502.05061. Bibcode:2014NatSR...4.6496S. doi:10.1038/srep06496. PMC 4178292. PMID 25263231. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4178292
Abdulla, N., et al. "Arabic sentiment analysis: Corpus-based and lexicon-based." Proceedings of the IEEE conference on Applied Electrical Engineering and Computing Technologies (AEECT). 2013.
Abooraig, Raddad; Al-Zu'bi, Shadi; Kanan, Tarek; Hawashin, Bilal; Al Ayoub, Mahmoud; Hmeidi, Ismail (June 2018). "Automatic categorization of Arabic articles based on their political orientation". Digital Investigation. 25: 24–41. doi:10.1016/j.diin.2018.04.003. /wiki/Doi_(identifier)
Kawala, François, et al. "Prédictions d'activité dans les réseaux sociaux en ligne." 4ième conférence sur les modèles et l'analyse des réseaux: Approches mathématiques et informatiques. 2013. https://hal.archives-ouvertes.fr/hal-00881395/document
Sabharwal, Ashish; Samulowitz, Horst; Tesauro, Gerald (2015). "Selecting Near-Optimal Learners via Incremental Data Allocation". arXiv:1601.00024 [cs.LG]. /wiki/ArXiv_(identifier)
Xu et al. "SemEval-2015 Task 1: Paraphrase and Semantic Similarity in Twitter (PIT)" Proceedings of the 9th International Workshop on Semantic Evaluation. 2015. https://www.aclweb.org/anthology/S15-2001
Xu et al. "Extracting Lexically Divergent Paraphrases from Twitter" Transactions of the Association for Computational (TACL). 2014. https://transacl.org/ojs/index.php/tacl/article/viewFile/498/64
Middleton, Stuart E; Middleton, Lee; Modafferi, Stefano (2014). "Real-Time Crisis Mapping of Natural Disasters Using Social Media" (PDF). IEEE Intelligent Systems. 29 (2): 9–17. doi:10.1109/MIS.2013.126. S2CID 15139204. https://eprints.soton.ac.uk/370581/1/ieee-is2014.pdf
"geoparsepy". 2016. Python PyPI library https://pypi.org/project/geoparsepy
Shmueli, Boaz; Ku, Lun-Wei; Ray, Soumya (2020). "Reactive Supervision: A New Method for Collecting Sarcasm Data". Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics. pp. 2553–2559. doi:10.18653/v1/2020.emnlp-main.201. S2CID 221970454. https://aclanthology.org/2020.emnlp-main.201/
Shmueli, Boaz. "SPIRS Sarcasm Dataset". GitHub. https://github.com/bshmueli/SPIRS
Gupta, Aakash (2020). "Dutch social media collection". COVID-19 Data Hub. doi:10.5072/FK2/MTPTL7. Retrieved 11 November 2023. https://huggingface.co/datasets/dutch_social/blob/main/dutch_social.py
"Streamlit". huggingface.co. Retrieved 18 December 2020. https://huggingface.co/datasets/viewer/?dataset=dutch_social
"Dutch Social media collection". kaggle.com. Retrieved 18 December 2020. https://kaggle.com/skylord/dutch-tweets
Shmueli, Boaz; Ray, Soumya; Lun-Wei (2021). "Happy Dance, Slow Clap: Using Reaction GIFs to Predict Induced Affect on Twitter". Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers). Vol. Association for Computational Linguistics. As. pp. 395–401. doi:10.18653/v1/2021.acl-short.50. S2CID 235125510. https://aclanthology.org/2021.acl-short.50
Shmueli, Boaz (5 May 2023), ReactionGIF, retrieved 6 October 2023 https://github.com/bshmueli/ReactionGIF
Forsyth, E., Lin, J., & Martell, C. (2008, June 25). The NPS Chat Corpus. Retrieved from http://faculty.nps.edu/cmartell/NPSChat.htm http://faculty.nps.edu/cmartell/NPSChat.htm
Sordoni, Alessandro; Galley, Michel; Auli, Michael; Brockett, Chris; Ji, Yangfeng; Mitchell, Margaret; Nie, Jian-Yun; Gao, Jianfeng; Dolan, Bill (2015). "A Neural Network Approach to Context-Sensitive Generation of Conversational Responses". arXiv:1506.06714 [cs.CL]. /wiki/ArXiv_(identifier)
Shaoul, C. & Westbury C. (2013) A reduced redundancy USENET corpus (2005–2011) Edmonton, AB: University of Alberta (downloaded from http://www.psych.ualberta.ca/~westburylab/downloads/usenetcorpus.download.html) http://www.psych.ualberta.ca/~westburylab/downloads/usenetcorpus.download.html
KAN, M. (2011, January). NUS Short Message Service (SMS) Corpus. Retrieved from http://www.comp.nus.edu.sg/entrepreneurship/innovation/osr/corpus/ Archived 29 June 2018 at the Wayback Machine http://www.comp.nus.edu.sg/entrepreneurship/innovation/osr/corpus/
Stuck_In_the_Matrix. (2015, July 3). I have every publicly available Reddit comment for research. ~ 1.7 billion comments @ 250 GB compressed. Any interest in this? [Original post]. Message posted to https://www.reddit.com/r/datasets/comments/3bxlg7/i_have_every_publicly_available_reddit_comment/ https://www.reddit.com/r/datasets/comments/3bxlg7/i_have_every_publicly_available_reddit_comment/
Lowe, Ryan; Pow, Nissan; Serban, Iulian; Pineau, Joelle (2015). "The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems". arXiv:1506.08909 [cs.CL]. /wiki/ArXiv_(identifier)
Jason Williams Antoine Raux Matthew Henderson, "[1]", Dialogue & Discourse | April 2016 . https://www.microsoft.com/en-us/research/publication/the-dialog-state-tracking-challenge-series-a-review/
Hoppe, Travis (16 December 2021), The-Pile-FreeLaw, retrieved 11 January 2023 https://github.com/thoppe/The-Pile-FreeLaw
Zheng, Lucia; Guha, Neel; Anderson, Brandon R.; Henderson, Peter; Ho, Daniel E. (21 June 2021). "When does pretraining help?". Proceedings of the Eighteenth International Conference on Artificial Intelligence and Law. New York, NY, USA: ACM. pp. 159–168. doi:10.1145/3462757.3466088. ISBN 9781450385268. S2CID 233296302. 9781450385268
"pile-of-law/pile-of-law · Datasets at Hugging Face". huggingface.co. 4 July 2022. Retrieved 11 January 2023. https://huggingface.co/datasets/pile-of-law/pile-of-law
"About | Caselaw Access Project". case.law. Retrieved 11 January 2023. https://case.law/about/
Roukos, Salim; Graff, David; Melamed, Dan (1995), Hansard French/English, Linguistic Data Consortium, doi:10.35111/JHGN-RV21, retrieved 26 February 2025 https://catalog.ldc.upenn.edu/LDC95T20
K. Kowsari, D. E. Brown, M. Heidarysafa, K. Jafari Meimandi, M. S. Gerber and L. E. Barnes, "HDLTex: Hierarchical Deep Learning for Text Classification", 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 364–371. doi:10.1109/ICMLA.2017.0-134 //doi.org/10.1109/ICMLA.2017.0-134
K. Kowsari, D. E. Brown, M. Heidarysafa, K. Jafari Meimandi, M. S. Gerber and L. E. Barnes, "Web of Science Dataset", doi:10.17632/9rw3vkcfy4.6 /wiki/Doi_(identifier)
Galgani, Filippo, Paul Compton, and Achim Hoffmann. "Combining different summarization techniques for legal text." Proceedings of the Workshop on Innovative Hybrid Approaches to the Processing of Textual Data. Association for Computational Linguistics, 2012. https://www.aclweb.org/anthology/W12-0515
Nagwani, N. K. (2015). "Summarizing large text collection using topic modeling and clustering based on MapReduce framework". Journal of Big Data. 2 (1): 1–18. doi:10.1186/s40537-015-0020-5. https://doi.org/10.1186%2Fs40537-015-0020-5
Schler, Jonathan; et al. (2006). "Effects of Age and Gender on Blogging" (PDF). AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs. 6. Archived from the original (PDF) on 14 November 2020. Retrieved 6 August 2019. https://web.archive.org/web/20201114000329/https://www.aaai.org/Papers/Symposia/Spring/2006/SS-06-03/SS06-03-039.pdf
Anand, Pranav, et al. "Believe Me-We Can Do This! Annotating Persuasive Acts in Blog Text."Computational Models of Natural Argument. 2011.
Traud, Amanda L., Peter J. Mucha, and Mason A. Porter. "Social structure of Facebook networks." Physica A: Statistical Mechanics and its Applications391.16 (2012): 4165–4180.
Richard, Emile; Savalle, Pierre-Andre; Vayatis, Nicolas (2012). "Estimation of Simultaneously Sparse and Low Rank Matrices". arXiv:1206.6474 [cs.DS]. /wiki/ArXiv_(identifier)
Richardson, Matthew; Burges, Christopher JC; Renshaw, Erin (2013). "MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text". EMNLP. 1. https://www.aclweb.org/anthology/D13-1020
Weston, Jason; Bordes, Antoine; Chopra, Sumit; Rush, Alexander M.; Bart van Merriënboer; Joulin, Armand; Mikolov, Tomas (2015). "Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks". arXiv:1502.05698 [cs.AI]. /wiki/ArXiv_(identifier)
Marcus, Mitchell P.; Ann Marcinkiewicz, Mary; Santorini, Beatrice (1993). "Building a large annotated corpus of English: The Penn Treebank". Computational Linguistics. 19 (2): 313–330. http://repository.upenn.edu/cgi/viewcontent.cgi?article=1246&context=cis_reports
Collins, Michael (2003). "Head-driven statistical models for natural language parsing". Computational Linguistics. 29 (4): 589–637. doi:10.1162/089120103322753356. https://doi.org/10.1162%2F089120103322753356
Guyon, Isabelle, et al., eds. Feature extraction: foundations and applications. Vol. 207. Springer, 2008. https://books.google.com/books?id=FOTzBwAAQBAJ&q=DEXTER
Lin, Yuri, et al. "Syntactic annotations for the google books ngram corpus." Proceedings of the ACL 2012 system demonstrations. Association for Computational Linguistics, 2012. https://www.aclweb.org/anthology/P/P12/P12-3029.pdf
Krishnamoorthy, Niveda; et al. (2013). "Generating Natural-Language Video Descriptions Using Text-Mined Knowledge". AAAI. 1. Archived from the original on 6 August 2019. Retrieved 6 August 2019. https://web.archive.org/web/20190806022756/https://www.aaai.org/ocs/index.php/AAAI/AAAI13/paper/download/6454/7204
Luyckx, Kim; Daelemans, Walter (2008). "Personae: a corpus for author and personality prediction from text". Proceedings of LREC-2008, the Sixth International Language Resources and Evaluation Conference. hdl:10067/687330151162165141. ISBN 978-2-9517408-4-6. 978-2-9517408-4-6
Solorio, Thamar, Ragib Hasan, and Mainul Mizan. "A case study of sockpuppet detection in wikipedia." Workshop on Language Analysis in Social Media (LASM) at NAACL HLT. 2013. https://www.aclweb.org/anthology/W13-1107
"Pushshift Files". files.pushshift.io. Archived from the original on 12 January 2023. Retrieved 12 January 2023. https://web.archive.org/web/20230112015822/https://files.pushshift.io/
Baumgartner, Jason; Zannettou, Savvas; Keegan, Brian; Squire, Megan; Blackburn, Jeremy (23 January 2020). "The Pushshift Reddit Dataset". arXiv:2001.08435 [cs.SI]. /wiki/ArXiv_(identifier)
Ciarelli, Patrick Marques; Oliveira, Elias (2009). "Agglomeration and Elimination of Terms for Dimensionality Reduction". 2009 Ninth International Conference on Intelligent Systems Design and Applications. pp. 547–552. doi:10.1109/ISDA.2009.9. ISBN 978-1-4244-4735-0. 978-1-4244-4735-0
Zhou, Mingyuan; Padilla, Oscar Hernan Madrid; Scott, James G. (2 July 2016). "Priors for Random Count Matrices Derived from a Family of Negative Binomial Processes". Journal of the American Statistical Association. 111 (515): 1144–1156. arXiv:1404.3331. doi:10.1080/01621459.2015.1075407. /wiki/ArXiv_(identifier)
Kotzias, Dimitrios, et al. "From group to individual labels using deep features." Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, 2015. http://datalab.ics.uci.edu/papers/kdd2015_dimitris.pdf
Ning, Yue; Muthiah, Sathappan; Rangwala, Huzefa; Ramakrishnan, Naren (2016). "Modeling Precursors for Event Forecasting via Nested Multi-Instance Learning". arXiv:1602.08033 [cs.SI]. /wiki/ArXiv_(identifier)
Buza, Krisztian. "Feedback prediction for blogs."Data analysis, machine learning and knowledge discovery. Springer International Publishing, 2014. 145–152. http://www.cs.bme.hu/~buza/pdfs/gfkl2012_blogs.pdf
Soysal, Ömer M (2015). "Association rule mining with mostly associated sequential patterns". Expert Systems with Applications. 42 (5): 2582–2592. doi:10.1016/j.eswa.2014.10.049. /wiki/Doi_(identifier)
Zhu, Yukun, et al. "Aligning books and movies: Towards story-like visual explanations by watching movies and reading books." Proceedings of the IEEE international conference on computer vision. 2015.
Bowman, Samuel R.; Angeli, Gabor; Potts, Christopher; Manning, Christopher D. (2015). "A large annotated corpus for learning natural language inference". arXiv:1508.05326 [cs.CL]. /wiki/ArXiv_(identifier)
"DSL Corpus Collection". ttg.uni-saarland.de. Retrieved 22 September 2017. http://ttg.uni-saarland.de/resources/DSLCC/
"Urban Dictionary Words and Definitions". https://www.kaggle.com/therohk/urban-dictionary-words-dataset
H. Elsahar, P. Vougiouklis, A. Remaci, C. Gravier, J. Hare, F. Laforest, E. Simperl, "T-REx: A Large Scale Alignment of Natural Language with Knowledge Base Triples", Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC-2018). https://www.aclweb.org/anthology/L18-1544
Wang, Alex; Singh, Amanpreet; Michael, Julian; Hill, Felix; Levy, Omer; Bowman, Samuel R. (2018). "GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding". arXiv:1804.07461 [cs.CL]. /wiki/ArXiv_(identifier)
"Computers Are Learning to Read—But They're Still Not So Smart". Wired. Retrieved 29 December 2019. https://www.wired.com/story/computers-are-learning-to-read-but-theyre-still-not-so-smart/
"GLUE Benchmark". gluebenchmark.com. Retrieved 25 February 2019. https://gluebenchmark.com/
Quan, Hoang Lam; Quang, Duy Le; Van Kiet, Nguyen; Ngan, Luu-Thuy Nguyen. "UIT-ViIC: A Dataset for the First Evaluation on Vietnamese Image Captioning". https://www.springerprofessional.de/uit-viic-a-dataset-for-the-first-evaluation-on-vietnamese-image-/18612672
To, Quoc Huy; Nguyen, Van Kiet; Nguyen, Luu Thuy Ngan; Nguyen, Gia Tuan Anh (2020). "Gender Prediction Based on Vietnamese Names with Machine Learning Techniques". Proceedings of the 4th International Conference on Natural Language Processing and Information Retrieval. pp. 55–60. arXiv:2010.10852. doi:10.1145/3443279.3443309. ISBN 9781450377607. S2CID 224814110. 9781450377607
Nguyen, Luan Thanh; Van Nguyen, Kiet; Nguyen, Ngan Luu-Thuy (18 March 2021). "Constructive and Toxic Speech Detection for Open-Domain Social Media Comments in Vietnamese". Advances and Trends in Artificial Intelligence. Artificial Intelligence Practices. Lecture Notes in Computer Science. Vol. 12798. pp. 572–583. arXiv:2103.10069. doi:10.1007/978-3-030-79457-6_49. ISBN 978-3-030-79456-9. S2CID 232269671. 978-3-030-79456-9
Saxton, David, et al. "Analysing Mathematical Reasoning Abilities of Neural Models." International Conference on Learning Representations. 2018.
Godfrey, J.J.; Holliman, E.C.; McDaniel, J. (1992). "SWITCHBOARD: Telephone speech corpus for research and development". [Proceedings] ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing. IEEE. pp. 517-520 vol.1. doi:10.1109/icassp.1992.225858. ISBN 0-7803-0532-9. 0-7803-0532-9
"Switchboard-1 Release 2 - Linguistic Data Consortium". catalog.ldc.upenn.edu. Retrieved 30 November 2024. https://catalog.ldc.upenn.edu/LDC97S62
Godfrey, J.J.; Holliman, E.C.; McDaniel, J. (1992). "SWITCHBOARD: Telephone speech corpus for research and development". [Proceedings] ICASSP-92: 1992 IEEE International Conference on Acoustics, Speech, and Signal Processing. IEEE. pp. 517-520 vol.1. doi:10.1109/icassp.1992.225858. ISBN 0-7803-0532-9. 0-7803-0532-9
"Switchboard-1 Release 2 - Linguistic Data Consortium". catalog.ldc.upenn.edu. Retrieved 30 November 2024. https://catalog.ldc.upenn.edu/LDC97S62
M. Versteegh, R. Thiollière, T. Schatz, X.-N. Cao, X. Anguera, A. Jansen, and E. Dupoux (2015). "The Zero Resource Speech Challenge 2015," in INTERSPEECH-2015.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, (2016). "The Zero Resource Speech Challenge 2015: Proposed Approaches and Results," in SLTU-2016. https://core.ac.uk/download/pdf/82574050.pdf
Sakar, Betul Erdogdu; et al. (2013). "Collection and analysis of a Parkinson speech dataset with multiple types of sound recordings". IEEE Journal of Biomedical and Health Informatics. 17 (4): 828–834. doi:10.1109/jbhi.2013.2245674. PMID 25055311. S2CID 15491516. /wiki/Doi_(identifier)
Zhao, Shunan; Rudzicz, Frank; Carvalho, Leonardo G.; Marquez-Chin, Cesar; Livingstone, Steven (2014). "Automatic detection of expressed emotion in Parkinson's Disease". 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 4813–4817. doi:10.1109/ICASSP.2014.6854516. ISBN 978-1-4799-2893-4. 978-1-4799-2893-4
Hammami, Nacereddine; Bedda, Mouldi (July 2010). "Improved tree model for arabic speech recognition". 2010 3rd International Conference on Computer Science and Information Technology. pp. 521–526. doi:10.1109/ICCSIT.2010.5563892. ISBN 978-1-4244-5537-9. 978-1-4244-5537-9
Maaten, Laurens. "Learning discriminative fisher kernels." Proceedings of the 28th International Conference on Machine Learning (ICML-11). 2011. https://lvdmaaten.github.io/publications/papers/ICML_2011.pdf
Cole, Ronald, and Mark Fanty. "Spoken letter recognition." Proc. Third DARPA Speech and Natural Language Workshop. 1990. https://www.aclweb.org/anthology/H90-1075
Chapelle, Olivier; Sindhwani, Vikas; Keerthi, Sathiya S. (2008). "Optimization techniques for semi-supervised support vector machines" (PDF). The Journal of Machine Learning Research. 9: 203–233. http://www.jmlr.org/papers/volume9/chapelle08a/chapelle08a.pdf
Kudo, Mineichi; Toyama, Jun; Shimbo, Masaru (November 1999). "Multidimensional curve classification using passing-through regions". Pattern Recognition Letters. 20 (11–13): 1103–1111. Bibcode:1999PaReL..20.1103K. doi:10.1016/s0167-8655(99)00077-x. /wiki/Bibcode_(identifier)
Jaeger, Herbert; Lukoševičius, Mantas; Popovici, Dan; Siewert, Udo (April 2007). "Optimization and applications of echo state networks with leaky- integrator neurons". Neural Networks. 20 (3): 335–352. doi:10.1016/j.neunet.2007.04.016. PMID 17517495. /wiki/Doi_(identifier)
Tsanas, A.; Little, M.A.; McSharry, P.E.; Ramig, L.O. (April 2010). "Accurate Telemonitoring of Parkinson's Disease Progression by Noninvasive Speech Tests". IEEE Transactions on Biomedical Engineering. 57 (4): 884–893. doi:10.1109/tbme.2009.2036000. PMID 19932995. /wiki/Doi_(identifier)
Clifford, Gari D.; Clifton, David (2012). "Wireless technology in disease management and medicine". Annual Review of Medicine. 63: 479–492. doi:10.1146/annurev-med-051210-114650. PMID 22053737. /wiki/Doi_(identifier)
Zue, Victor; Seneff, Stephanie; Glass, James (1990). "Speech database development at MIT: TIMIT and beyond". Speech Communication. 9 (4): 351–356. doi:10.1016/0167-6393(90)90010-7. /wiki/Doi_(identifier)
Kapadia, S.; Valtchev, V.; Young, S.J. (1993). "MMI training for continuous phoneme recognition on the TIMIT database". IEEE International Conference on Acoustics Speech and Signal Processing. pp. 491-494 vol.2. doi:10.1109/ICASSP.1993.319349. ISBN 0-7803-0946-4. 0-7803-0946-4
Halabi, Nawar (2016). Modern Standard Arabic Phonetics for Speech Synthesis (PDF) (PhD Thesis). University of Southampton, School of Electronics and Computer Science. http://en.arabicspeechcorpus.com/Nawar%20Halabi%20PhD%20Thesis%20Revised.pdf
Ardila, Rosana; Branson, Megan; Davis, Kelly; Henretty, Michael; Kohler, Michael; Meyer, Josh; Morais, Reuben; Saunders, Lindsay; Tyers, Francis M.; Weber, Gregor (13 December 2019). "Common Voice: A Massively-Multilingual Speech Corpus". arXiv:1912.06670v2 [cs.CL]. /wiki/ArXiv_(identifier)
"The LJ Speech Dataset". keithito.com. Retrieved 13 April 2022. https://keithito.com/LJ-Speech-Dataset
Ghandoura, Abdulkader; Hjabo, Farouk; Al Dakkak, Oumayma (June 2021). "Building and benchmarking an Arabic Speech Commands dataset for small-footprint keyword spotting". Engineering Applications of Artificial Intelligence. 102: 104267. doi:10.1016/j.engappai.2021.104267. /wiki/Doi_(identifier)
Zhou, Fang; Claire, Q.; King, Ross D. (2014). "Predicting the Geographical Origin of Music". 2014 IEEE International Conference on Data Mining. pp. 1115–1120. doi:10.1109/ICDM.2014.73. ISBN 978-1-4799-4302-9. 978-1-4799-4302-9
Saccenti, Edoardo; Camacho, José (2015). "On the use of the observation-wise k-fold operation in PCA cross-validation". Journal of Chemometrics. 29 (8): 467–478. doi:10.1002/cem.2726. hdl:10481/55302. S2CID 62248957. /wiki/Doi_(identifier)
Bertin-Mahieux, Thierry, et al. "The million song dataset." ISMIR 2011: Proceedings of the 12th International Society for Music Information Retrieval Conference, 24–28 October 2011, Miami, Florida. University of Miami, 2011.
Henaff, Mikael; et al. (2011). "Unsupervised learning of sparse features for scalable audio classification" (PDF). ISMIR. 11. https://archives.ismir.net/ismir2011/paper/000128.pdf
Rafii, Zafar (2017). "Music". MUSDB18 – a corpus for music separation. doi:10.5281/zenodo.1117372. /wiki/Doi_(identifier)
Defferrard, Michaël; Benzi, Kirell; Vandergheynst, Pierre; Bresson, Xavier (6 December 2016). "FMA: A Dataset For Music Analysis". arXiv:1612.01840 [cs.SD]. /wiki/ArXiv_(identifier)
Esposito, Roberto; Radicioni, Daniele P. (2009). "Carpediem: Optimizing the viterbi algorithm and applications to supervised sequential learning" (PDF). The Journal of Machine Learning Research. 10: 1851–1880. http://www.jmlr.org/papers/volume10/esposito09a/esposito09a.pdf
Sourati, Jamshid; et al. (2016). "Classification Active Learning Based on Mutual Information". Entropy. 18 (2): 51. Bibcode:2016Entrp..18...51S. doi:10.3390/e18020051. https://doi.org/10.3390%2Fe18020051
Salamon, Justin; Jacoby, Christopher; Bello, Juan Pablo. "A dataset and taxonomy for urban sound research." Proceedings of the ACM International Conference on Multimedia. ACM, 2014. https://www.researchgate.net/profile/Justin_Salamon/publication/267269056_A_Dataset_and_Taxonomy_for_Urban_Sound_Research/links/544936af0cf2f63880810a84/A-Dataset-and-Taxonomy-for-Urban-Sound-Research.pdf
Lagrange, Mathieu; Lafay, Grégoire; Rossignol, Mathias; Benetos, Emmanouil; Roebel, Axel (2015). "An evaluation framework for event detection using a morphological model of acoustic scenes". arXiv:1502.00141 [stat.ML]. /wiki/ArXiv_(identifier)
Gemmeke, Jort F., et al. "Audio Set: An ontology and human-labeled dataset for audio events." IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). 2017. /wiki/IEEE
"Watch out, birders: Artificial intelligence has learned to spot birds from their songs". Science | AAAS. 18 July 2018. Retrieved 22 July 2018. https://www.science.org/content/article/watch-out-birders-artificial-intelligence-has-learned-spot-birds-their-songs
"Bird Audio Detection challenge". Machine Listening Lab at Queen Mary University. 3 May 2016. Retrieved 22 July 2018. http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/
Wichern, Gordon; Antognini, Joe; Flynn, Michael; Licheng Richard Zhu; McQuinn, Emmett; Crow, Dwight; Manilow, Ethan; Jonathan Le Roux (2019). "WHAM!: Extending Speech Separation to Noisy Environments". arXiv:1907.01160 [cs.SD]. /wiki/ArXiv_(identifier)
Drossos, K., Lipping, S., and Virtanen, T. "Clotho: An Audio Captioning Dataset" IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). 2020. /wiki/IEEE
Drossos, K., Lipping, S., and Virtanen, T. (2019). Clotho dataset (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3490684 /wiki/Zenodo
The CAIDA UCSD Dataset on the Witty Worm – 19–24 March 2004, http://www.caida.org/data/passive/witty_worm_dataset.xml http://www.caida.org/data/passive/witty_worm_dataset.xml
Chen, Zesheng; Ji, Chuanyi (2007). "Optimal worm-scanning method using vulnerable-host distributions". International Journal of Security and Networks. 2 (1/2): 71. doi:10.1504/IJSN.2007.012826. /wiki/Doi_(identifier)
Kachuee, Mohamad; Kiani, Mohammad Mahdi; Mohammadzade, Hoda; Shabany, Mahdi (2015). "Cuff-less high-accuracy calibration-free blood pressure estimation using pulse transit time". 2015 IEEE International Symposium on Circuits and Systems (ISCAS). pp. 1006–1009. doi:10.1109/ISCAS.2015.7168806. ISBN 978-1-4799-8391-9. 978-1-4799-8391-9
Goldberger, Ary L.; Amaral, Luis A. N.; Glass, Leon; Hausdorff, Jeffrey M.; Ivanov, Plamen Ch.; Mark, Roger G.; Mietus, Joseph E.; Moody, George B.; Peng, Chung-Kang; Stanley, H. Eugene (13 June 2000). "PhysioBank, PhysioToolkit, and PhysioNet: Components of a New Research Resource for Complex Physiologic Signals". Circulation. 101 (23): E215-20. doi:10.1161/01.CIR.101.23.e215. PMID 10851218. /wiki/Doi_(identifier)
Vergara, Alexander; et al. (2012). "Chemical gas sensor drift compensation using classifier ensembles". Sensors and Actuators B: Chemical. 166: 320–329. Bibcode:2012SeAcB.166..320V. doi:10.1016/j.snb.2012.01.074. /wiki/Bibcode_(identifier)
Korotcenkov, G.; Cho, B. K. (2014). "Engineering approaches to improvement of conductometric gas sensor parameters. Part 2: Decrease of dissipated (consumable) power and improvement stability and reliability". Sensors and Actuators B: Chemical. 198: 316–341. Bibcode:2014SeAcB.198..316K. doi:10.1016/j.snb.2014.03.069. /wiki/Bibcode_(identifier)
Quinlan, John R (1992). "Learning with continuous classes" (PDF). 5th Australian Joint Conference on Artificial Intelligence. 92. https://sci2s.ugr.es/keel/pdf/algorithm/congreso/1992-Quinlan-AI.pdf
Merz, Christopher J.; Pazzani, Michael J. (1999). "A principal components approach to combining regression estimates". Machine Learning. 36 (1–2): 9–32. doi:10.1023/a:1007507221352. https://doi.org/10.1023%2Fa%3A1007507221352
Torres-Sospedra, Joaquin, et al. "UJIIndoorLoc-Mag: A new database for magnetic field-based localization problems." Indoor Positioning and Indoor Navigation (IPIN), 2015 International Conference on. IEEE, 2015.
Berkvens, Rafael, Maarten Weyn, and Herbert Peremans. "Mean Mutual Information of Probabilistic Wi-Fi Localization." Indoor Positioning and Indoor Navigation (IPIN), 2015 International Conference on. Banff, Canada: IPIN. 2015. https://www.researchgate.net/profile/Raf_Berkvens/publication/284154212_Mean_Mutual_Information_of_Probabilistic_Wi-Fi_Localization/links/564c6b7508aeab8ed5e92fcb.pdf
Paschke, Fabian, et al. "Sensorlose Zustandsüberwachung an Synchronmotoren."Proceedings. 23. Workshop Computational Intelligence, Dortmund, 5.-6. Dezember 2013. KIT Scientific Publishing, 2013.
Lessmeier, Christian, et al. "Data Acquisition and Signal Analysis from Measured Motor Currents for Defect Detection in Electromechanical Drive Systems." https://www.researchgate.net/profile/Olaf_Enge-Rosenblatt/publication/264441239_Data_Acquisition_and_Signal_Analysis_from_Measured_Motor_Currents_for_Defect_Detection_in_Electromechanical_Drive_Systems/links/53df97e90cf2a768e49bb3b9.pdf
Ugulino, Wallace, et al. "Wearable computing: Accelerometers’ data classification of body postures and movements Archived 25 September 2020 at the Wayback Machine." Advances in Artificial Intelligence-SBIA 2012. Springer Berlin Heidelberg, 2012. 52–61. http://groupware.secondlab.inf.puc-rio.br/public/papers/2012.Ugulino.WearableComputing.HAR.Classifier.RIBBON.pdf
Schneider, Jan; et al. (2015). "Augmenting the senses: a review on sensor-based learning support". Sensors. 15 (2): 4097–4133. Bibcode:2015Senso..15.4097S. doi:10.3390/s150204097. PMC 4367401. PMID 25679313. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4367401
Madeo, Renata CB, Clodoaldo AM Lima, and Sarajane M. Peres. "Gesture unit segmentation using support vector machines: segmenting gestures from rest positions." Proceedings of the 28th Annual ACM Symposium on Applied Computing. ACM, 2013. https://tarjomefa.com/wp-content/uploads/2016/11/5781-English.pdf
Lun, Roanna; Zhao, Wenbing (2015). "A survey of applications and human motion recognition with Microsoft Kinect". International Journal of Pattern Recognition and Artificial Intelligence. 29 (5): 1555008. doi:10.1142/s0218001415550083. https://engagedscholarship.csuohio.edu/cgi/viewcontent.cgi?article=1417&context=enece_facpub
Theodoridis, Theodoros; Huosheng Hu (2007). "Action classification of 3D human models using dynamic ANNs for mobile robot surveillance". 2007 IEEE International Conference on Robotics and Biomimetics (ROBIO). pp. 371–376. doi:10.1109/ROBIO.2007.4522190. ISBN 978-1-4244-1761-2. 978-1-4244-1761-2
Etemad, Seyed Ali; Arya, Ali (2009). "3D human action recognition and style transformation using resilient backpropagation neural networks". 2009 IEEE International Conference on Intelligent Computing and Intelligent Systems. pp. 296–301. doi:10.1109/ICICISYS.2009.5357690. ISBN 978-1-4244-4754-1. 978-1-4244-4754-1
Altun, Kerem; Barshan, Billur; Tunçel, Orkun (2010). "Comparative study on classifying human activities with miniature inertial and magnetic sensors". Pattern Recognition. 43 (10): 3605–3620. Bibcode:2010PatRe..43.3605A. doi:10.1016/j.patcog.2010.04.019. hdl:11693/11947. /wiki/Bibcode_(identifier)
Nathan, Ran; et al. (2012). "Using tri-axial acceleration data to identify behavioral modes of free-ranging animals: general concepts and tools illustrated for griffon vultures". The Journal of Experimental Biology. 215 (6): 986–996. Bibcode:2012JExpB.215..986N. doi:10.1242/jeb.058602. PMC 3284320. PMID 22357592. /wiki/Ran_Nathan
Anguita, Davide, et al. "Human activity recognition on smartphones using a multiclass hardware-friendly support vector machine." Ambient assisted living and home care. Springer Berlin Heidelberg, 2012. 216–223. https://upcommons.upc.edu/bitstream/handle/2117/101769/IWAAL2012.pdf
Su, Xing; Tong, Hanghang; Ji, Ping (2014). "Activity recognition with smartphone sensors". Tsinghua Science and Technology. 19 (3): 235–249. doi:10.1109/tst.2014.6838194. S2CID 62751498. /wiki/Doi_(identifier)
Kadous, Mohammed Waleed. Temporal classification: Extending the classification paradigm to multivariate time series. Diss. The University of New South Wales, 2002. https://pdfs.semanticscholar.org/4bad/c3f0ad169ed9ec7d073375e9b168fa9f6c8f.pdf
Graves, Alex, et al. "Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks." Proceedings of the 23rd international conference on Machine learning. ACM, 2006. https://mediatum.ub.tum.de/doc/1292048/file.pdf
Velloso, Eduardo, et al. "Qualitative activity recognition of weight lifting exercises."Proceedings of the 4th Augmented Human International Conference. ACM, 2013. https://www.perceptualui.org/publications/velloso13_ah.pdf
Mortazavi, Bobak Jack, et al. "Determining the single best axis for exercise repetition recognition and counting on smartwatches Archived 4 November 2021 at the Wayback Machine." Wearable and Implantable Body Sensor Networks (BSN), 2014 11th International Conference on. IEEE, 2014. http://www.thehabitslab.com/assets/papers/28.pdf
Sapsanis, Christos, et al. "Improving EMG based Classification of basic hand movements using EMD." Engineering in Medicine and Biology Society (EMBC), 2013 35th Annual International Conference of the IEEE. IEEE, 2013. https://www.researchgate.net/profile/Christos_Sapsanis/publication/257602303_Improving_EMG_based_classification_of_basic_hand_movements_using_EMD/links/56dfb7fd08ae979addef64a2/Improving-EMG-based-classification-of-basic-hand-movements-using-EMD.pdf
Andrianesis, Konstantinos; Tzes, Anthony (2015). "Development and control of a multifunctional prosthetic hand with shape memory alloy actuators". Journal of Intelligent & Robotic Systems. 78 (2): 257–289. doi:10.1007/s10846-014-0061-6. S2CID 207174078. /wiki/Doi_(identifier)
Andrianesis, Konstantinos; Tzes, Anthony (2015). "Development and control of a multifunctional prosthetic hand with shape memory alloy actuators". Journal of Intelligent & Robotic Systems. 78 (2): 257–289. doi:10.1007/s10846-014-0061-6. S2CID 207174078. /wiki/Doi_(identifier)
Banos, Oresti; et al. (2014). "Dealing with the effects of sensor displacement in wearable activity recognition". Sensors. 14 (6): 9995–10023. Bibcode:2014Senso..14.9995B. doi:10.3390/s140609995. PMC 4118358. PMID 24915181. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4118358
Stisen, Allan; Blunck, Henrik; Bhattacharya, Sourav; Prentow, Thor Siiger; Kjærgaard, Mikkel Baun; Dey, Anind; Sonne, Tobias; Jensen, Mads Møller (2015). "Smart Devices are Different: Assessing and MitigatingMobile Sensing Heterogeneities for Activity Recognition". Proceedings of the 13th ACM Conference on Embedded Networked Sensor Systems. pp. 127–140. doi:10.1145/2809695.2809718. ISBN 978-1-4503-3631-4. 978-1-4503-3631-4
Bhattacharya, Sourav; Lane, Nicholas D. (2016). "From smart to deep: Robust activity recognition on smartwatches using deep learning". 2016 IEEE International Conference on Pervasive Computing and Communication Workshops (PerCom Workshops). pp. 1–6. doi:10.1109/PERCOMW.2016.7457169. ISBN 978-1-5090-1941-0. 978-1-5090-1941-0
Bacciu, Davide; et al. (2014). "An experimental characterization of reservoir computing in ambient assisted living applications". Neural Computing and Applications. 24 (6): 1451–1464. doi:10.1007/s00521-013-1364-4. hdl:11568/237959. S2CID 14124013. /wiki/Doi_(identifier)
Palumbo, Filippo; Barsocchi, Paolo; Gallicchio, Claudio; Chessa, Stefano; Micheli, Alessio (2013). "Multisensor Data Fusion for Activity Recognition Based on Reservoir Computing". Evaluating AAL Systems Through Competitive Benchmarking. Communications in Computer and Information Science. Vol. 386. pp. 24–35. doi:10.1007/978-3-642-41043-7_3. ISBN 978-3-642-41042-0. 978-3-642-41042-0
Reiss, Attila; Stricker, Didier (2012). "Introducing a New Benchmarked Dataset for Activity Monitoring". 2012 16th International Symposium on Wearable Computers. pp. 108–109. doi:10.1109/ISWC.2012.13. ISBN 978-0-7695-4697-1. 978-0-7695-4697-1
Roggen, Daniel; Forster, Kilian; Calatroni, Alberto; Holleczek, Thomas; Fang, Yu; Troster, Gerhard; Ferscha, Alois; Holzmann, Clemens; Riener, Andreas; Lukowicz, Paul; Pirkl, Gerald; Bannach, David; Kunze, Kai; Chavarriaga, Ricardo; Millan, Jose del R. (2009). "OPPORTUNITY: Towards opportunistic activity and context recognition systems". 2009 IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks & Workshops. pp. 1–6. doi:10.1109/WOWMOM.2009.5282442. ISBN 978-1-4244-4440-3. 978-1-4244-4440-3
Kurz, Marc, et al. "Dynamic quantification of activity recognition capabilities in opportunistic systems." Vehicular Technology Conference (VTC Spring), 2011 IEEE 73rd. IEEE, 2011. https://www.researchgate.net/profile/Marc_Kurz/publication/220271166_Dynamic_Quantification_of_Activity_Recognition_Capabilities_in_Opportunistic_Systems/links/09e4150f66b480c97a000000/Dynamic-Quantification-of-Activity-Recognition-Capabilities-in-Opportunistic-Systems.pdf
Sztyler, Timo; Stuckenschmidt, Heiner (2016). "On-body localization of wearable devices: An investigation of position-aware activity recognition". 2016 IEEE International Conference on Pervasive Computing and Communications (PerCom). pp. 1–9. doi:10.1109/PERCOM.2016.7456521. ISBN 978-1-4673-8779-8. 978-1-4673-8779-8
Zhi, Ying Xuan; Lukasik, Michelle; Li, Michael H.; Dolatabadi, Elham; Wang, Rosalie H.; Taati, Babak (2018). "Automatic Detection of Compensation During Robotic Stroke Rehabilitation Therapy". IEEE Journal of Translational Engineering in Health and Medicine. 6: 1–7. doi:10.1109/JTEHM.2017.2780836. PMC 5788403. PMID 29404226. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5788403
Dolatabadi, Elham; Zhi, Ying Xuan; Ye, Bing; Coahran, Marge; Lupinacci, Giorgia; Mihailidis, Alex; Wang, Rosalie; Taati, Babak (2017). "The toronto rehab stroke pose dataset to detect compensation during stroke rehabilitation therapy". Proceedings of the 11th EAI International Conference on Pervasive Computing Technologies for Healthcare. pp. 375–381. doi:10.1145/3154862.3154925. ISBN 978-1-4503-6363-1. 978-1-4503-6363-1
"Toronto Rehab Stroke Pose Dataset". https://www.kaggle.com/derekdb/toronto-robot-stroke-posture-dataset
Jung, Merel M.; Poel, Mannes; Poppe, Ronald; Heylen, Dirk K. J. (March 2017). "Automatic recognition of touch gestures in the corpus of social touch". Journal on Multimodal User Interfaces. 11 (1): 81–96. doi:10.1007/s12193-016-0232-9. /wiki/Doi_(identifier)
Jung, M.M. (Merel) (1 June 2016). "Corpus of Social Touch (CoST)". University of Twente. doi:10.4121/uuid:5ef62345-3b3e-479c-8e1d-c922748c9b29. {{cite journal}}: Cite journal requires |journal= (help) https://data.4tu.nl/articles/dataset/Corpus_of_Social_Touch_CoST_/12696869
Aeberhard, S., D. Coomans, and O. De Vel. "Comparison of classifiers in high dimensional settings." Dept. Math. Statist., James Cook Univ., North Queensland, Australia, Tech. Rep 92-02 (1992).
Basu, Sugato. "Semi-supervised clustering with limited background knowledge." AAAI. 2004. http://www.aaai.org/Papers/AAAI/2004/AAAI04-138.pdf
Tüfekci, Pınar (2014). "Prediction of full load electrical power output of a base load operated combined cycle power plant using machine learning methods". International Journal of Electrical Power & Energy Systems. 60: 126–140. Bibcode:2014IJEPE..60..126T. doi:10.1016/j.ijepes.2014.02.027. https://hal.science/hal-04823875
Kaya, Heysem, Pınar Tüfekci, and Fikret S. Gürgen. "Local and global learning methods for predicting power of a combined gas & steam turbine." International conference on emerging trends in computer and electronics engineering (ICETCEE'2012), Dubai. 2012.
Baldi, Pierre; Sadowski, Peter; Whiteson, Daniel (2014). "Searching for exotic particles in high-energy physics with deep learning". Nature Communications. 5: 2014. arXiv:1402.4735. Bibcode:2014NatCo...5.4308B. doi:10.1038/ncomms5308. PMID 24986233. S2CID 195953. /wiki/ArXiv_(identifier)
Baldi, Pierre; Sadowski, Peter; Whiteson, Daniel (2015). "Enhanced Higgs Boson to τ+ τ− Search with Deep Learning". Physical Review Letters. 114 (11): 111801. arXiv:1410.3469. Bibcode:2015PhRvL.114k1801B. doi:10.1103/physrevlett.114.111801. PMID 25839260. S2CID 2339142. /wiki/ArXiv_(identifier)
Adam-Bourdarios, C.; Cowan, G.; Germain-Renaud, C.; Guyon, I.; Kégl, B.; Rousseau, D. (2015). "The Higgs Machine Learning Challenge". Journal of Physics: Conference Series. 664 (7): 072015. Bibcode:2015JPhCS.664g2015A. doi:10.1088/1742-6596/664/7/072015. https://higgsml.lal.in2p3.fr/
Baldi, Pierre; Sadowski, Peter; Whiteson, Daniel (2015). "Enhanced Higgs Boson to τ+ τ− Search with Deep Learning". Physical Review Letters. 114 (11): 111801. arXiv:1410.3469. Bibcode:2015PhRvL.114k1801B. doi:10.1103/physrevlett.114.111801. PMID 25839260. S2CID 2339142. /wiki/ArXiv_(identifier)
Adam-Bourdarios, C.; Cowan, G.; Germain-Renaud, C.; Guyon, I.; Kégl, B.; Rousseau, D. (2015). "The Higgs Machine Learning Challenge". Journal of Physics: Conference Series. 664 (7): 072015. Bibcode:2015JPhCS.664g2015A. doi:10.1088/1742-6596/664/7/072015. https://higgsml.lal.in2p3.fr/
Baldi, Pierre; Cranmer, Kyle; Faucett, Taylor; Sadowski, Peter; Whiteson, Daniel (2016). "Parameterized neural networks for high-energy physics". The European Physical Journal C. 76 (5): 235. arXiv:1601.07913. Bibcode:2016EPJC...76..235B. doi:10.1140/epjc/s10052-016-4099-4. S2CID 254108545. /wiki/ArXiv_(identifier)
Ortigosa, I.; Lopez, R.; Garcia, J. "A neural networks approach to residuary resistance of sailing yachts prediction". Proceedings of the International Conference on Marine Engineering MARINE. 2007.
Gerritsma, J., R. Onnink, and A. Versluis.Geometry, resistance and stability of the delft systematic yacht hull series. Delft University of Technology, 1981.
Liu, Huan, and Hiroshi Motoda. Feature extraction, construction and selection: A data mining perspective. Springer Science & Business Media, 1998. https://books.google.com/books?id=zi_0EdWW5fYC
Reich, Yoram. Converging to Ideal Design Knowledge by Learning. [Carnegie Mellon University], Engineering Design Research Center, 1989.
Todorovski, Ljupčo; Džeroski, Sašo (1999). "Experiments in Meta-level Learning with ILP". Principles of Data Mining and Knowledge Discovery. Lecture Notes in Computer Science. Vol. 1704. pp. 98–106. doi:10.1007/978-3-540-48247-5_11. ISBN 978-3-540-66490-1. S2CID 39382993. 978-3-540-66490-1
Wang, Yong. A new approach to fitting linear models in high dimensional spaces. Diss. The University of Waikato, 2000. http://www.cs.waikato.ac.nz/~ml/publications/2000/thesis.pdf
Kibler, Dennis; Aha, David W.; Albert, Marc K. (1989). "Instance-based prediction of real-valued attributes". Computational Intelligence. 5 (2): 51–57. doi:10.1111/j.1467-8640.1989.tb00315.x. S2CID 40800413. https://escholarship.org/uc/item/68f860zb
Palmer, Christopher R.; Faloutsos, Christos (2003). "Electricity Based External Similarity of Categorical Attributes". Advances in Knowledge Discovery and Data Mining. Lecture Notes in Computer Science. Vol. 2637. pp. 486–500. doi:10.1007/3-540-36175-8_49. ISBN 978-3-540-04760-5. 978-3-540-04760-5
Tsanas, Athanasios; Xifara, Angeliki (2012). "Accurate quantitative estimation of energy performance of residential buildings using statistical machine learning tools". Energy and Buildings. 49: 560–567. Bibcode:2012EneBu..49..560T. doi:10.1016/j.enbuild.2012.03.003. /wiki/Bibcode_(identifier)
De Wilde, Pieter (2014). "The gap between predicted and measured energy performance of buildings: A framework for investigation". Automation in Construction. 41: 40–49. doi:10.1016/j.autcon.2014.02.009. /wiki/Doi_(identifier)
Brooks, Thomas F., D. Stuart Pope, and Michael A. Marcolini. Airfoil self-noise and prediction. Vol. 1218. National Aeronautics and Space Administration, Office of Management, Scientific and Technical Information Division, 1989. https://ntrs.nasa.gov/archive/nasa/casi.ntrs.nasa.gov/19890016302.pdf
Draper, David. "Assessment and propagation of model uncertainty." Journal of the Royal Statistical Society, Series B (Methodological) (1995): 45–97. http://www2.denizyuret.com/ref/draper/assessment-and-propagation.pdf
Lavine, Michael (1991). "Problems in extrapolation illustrated with space shuttle O-ring data". Journal of the American Statistical Association. 86 (416): 919–921. doi:10.1080/01621459.1991.10475132. /wiki/Doi_(identifier)
Wang, J.; Yu, B.; Gasser, L. (2002). "Concept tree based clustering visualization with shaded similarity matrices". 2002 IEEE International Conference on Data Mining, 2002. Proceedings. pp. 697–700. doi:10.1109/ICDM.2002.1184032. ISBN 0-7695-1754-4. 0-7695-1754-4
Pettengill, Gordon H.; Ford, Peter G.; Johnson, William T. K.; Raney, R. Keith; Soderblom, Laurence A. (12 April 1991). "Magellan: Radar Performance and Data Products". Science. 252 (5003): 260–265. Bibcode:1991Sci...252..260P. doi:10.1126/science.252.5003.260. PMID 17769272. /wiki/Bibcode_(identifier)
Aharonian, F.; et al. (2008). "Energy spectrum of cosmic-ray electrons at TeV energies". Physical Review Letters. 101 (26): 261104. arXiv:0811.3894. Bibcode:2008PhRvL.101z1104A. doi:10.1103/PhysRevLett.101.261104. hdl:2440/51450. PMID 19437632. S2CID 41850528. /wiki/ArXiv_(identifier)
Aharonian, F.; et al. (2008). "Energy spectrum of cosmic-ray electrons at TeV energies". Physical Review Letters. 101 (26): 261104. arXiv:0811.3894. Bibcode:2008PhRvL.101z1104A. doi:10.1103/PhysRevLett.101.261104. hdl:2440/51450. PMID 19437632. S2CID 41850528. /wiki/ArXiv_(identifier)
Bock, R. K.; et al. (2004). "Methods for multidimensional event classification: a case study using images from a Cherenkov gamma-ray telescope". Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment. 516 (2): 511–528. Bibcode:2004NIMPA.516..511B. doi:10.1016/j.nima.2003.08.157. /wiki/Bibcode_(identifier)
Li, Jinyan; et al. (2004). "Deeps: A new instance-based lazy discovery and classification system". Machine Learning. 54 (2): 99–124. doi:10.1023/b:mach.0000011804.08528.7d. https://doi.org/10.1023%2Fb%3Amach.0000011804.08528.7d
Villaescusa-Navarro, Francisco; al., et (2022). "The CAMELS Multifield Data Set: Learning the Universe's Fundamental Parameters with Artificial Intelligence". The Astrophysical Journal Supplement Series. 259 (2): 61. arXiv:2109.10915. Bibcode:2022ApJS..259...61V. doi:10.3847/1538-4365/ac5ab0. S2CID 237604997. https://doi.org/10.3847%2F1538-4365%2Fac5ab0
Siebert, Lee, and Tom Simkin. "Volcanoes of the world: an illustrated catalog of Holocene volcanoes and their eruptions." (2014).
Sikora, Marek; Wróbel, Łukasz (2010). "Application of rule induction algorithms for analysis of data collected by seismic hazard monitoring systems in coal mines". Archives of Mining Sciences. 55 (1): 91–114. https://www.infona.pl/resource/bwmeta1.element.baztech-article-BPZ5-0008-0008
Sikora, Marek; Sikora, Beata (2012). "Rough Natural Hazards Monitoring". Rough Sets: Selected Methods and Applications in Management and Engineering. Advanced Information and Knowledge Processing. pp. 163–179. doi:10.1007/978-1-4471-2760-4_10. ISBN 978-1-4471-2759-8. 978-1-4471-2759-8
Addor, Nans; Newman, Andrew J.; Mizukami, Naoki; Clark, Martyn P. (20 October 2017). "The CAMELS data set: catchment attributes and meteorology for large-sample studies". Hydrology and Earth System Sciences. 21 (10): 5293–5313. Bibcode:2017HESS...21.5293A. doi:10.5194/hess-21-5293-2017. https://doi.org/10.5194%2Fhess-21-5293-2017
Newman, A. J.; Clark, M. P.; Sampson, K.; Wood, A.; Hay, L. E.; Bock, A.; Viger, R. J.; Blodgett, D.; Brekke, L.; Arnold, J. R.; Hopson, T.; Duan, Q. (14 January 2015). "Development of a large-sample watershed-scale hydrometeorological data set for the contiguous USA: data set characteristics and assessment of regional variability in hydrologic model performance". Hydrology and Earth System Sciences. 19 (1): 209–223. Bibcode:2015HESS...19..209N. doi:10.5194/hess-19-209-2015. https://doi.org/10.5194%2Fhess-19-209-2015
Alvarez-Garreton, Camila; Mendoza, Pablo A.; Boisier, Juan Pablo; Addor, Nans; Galleguillos, Mauricio; Zambrano-Bigiarini, Mauricio; Lara, Antonio; Puelma, Cristóbal; Cortes, Gonzalo; Garreaud, Rene; McPhee, James; Ayala, Alvaro (13 November 2018). "The CAMELS-CL dataset: catchment attributes and meteorology for large sample studies – Chile dataset". Hydrology and Earth System Sciences. 22 (11): 5817–5846. Bibcode:2018HESS...22.5817A. doi:10.5194/hess-22-5817-2018. https://doi.org/10.5194%2Fhess-22-5817-2018
Chagas, Vinícius B. P.; Chaffe, Pedro L. B.; Addor, Nans; Fan, Fernando M.; Fleischmann, Ayan S.; Paiva, Rodrigo C. D.; Siqueira, Vinícius A. (8 September 2020). "CAMELS-BR: hydrometeorological time series and landscape attributes for 897 catchments in Brazil". Earth System Science Data. 12 (3): 2075–2096. Bibcode:2020ESSD...12.2075C. doi:10.5194/essd-12-2075-2020. https://doi.org/10.5194%2Fessd-12-2075-2020
Coxon, Gemma; Addor, Nans; Bloomfield, John P.; Freer, Jim; Fry, Matt; Hannaford, Jamie; Howden, Nicholas J. K.; Lane, Rosanna; Lewis, Melinda; Robinson, Emma L.; Wagener, Thorsten; Woods, Ross (12 October 2020). "CAMELS-GB: hydrometeorological time series and landscape attributes for 671 catchments in Great Britain". Earth System Science Data. 12 (4): 2459–2483. Bibcode:2020ESSD...12.2459C. doi:10.5194/essd-12-2459-2020. https://doi.org/10.5194%2Fessd-12-2459-2020
Fowler, Keirnan J. A.; Acharya, Suwash Chandra; Addor, Nans; Chou, Chihchung; Peel, Murray C. (6 August 2021). "CAMELS-AUS: hydrometeorological time series and landscape attributes for 222 catchments in Australia". Earth System Science Data. 13 (8): 3847–3867. Bibcode:2021ESSD...13.3847F. doi:10.5194/essd-13-3847-2021. https://doi.org/10.5194%2Fessd-13-3847-2021
Klingler, Christoph; Schulz, Karsten; Herrnegger, Mathew (16 September 2021). "LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe". Earth System Science Data. 13 (9): 4529–4565. Bibcode:2021ESSD...13.4529K. doi:10.5194/essd-13-4529-2021. https://doi.org/10.5194%2Fessd-13-4529-2021
Yeh, I–C (1998). "Modeling of strength of high-performance concrete using artificial neural networks". Cement and Concrete Research. 28 (12): 1797–1808. doi:10.1016/s0008-8846(98)00165-3. /wiki/Doi_(identifier)
Zarandi, MH Fazel; et al. (2008). "Fuzzy polynomial neural networks for approximation of the compressive strength of concrete". Applied Soft Computing. 8 (1): 488–498. Bibcode:2008ApSoC...8...79S. doi:10.1016/j.asoc.2007.02.010. /wiki/Bibcode_(identifier)
Yeh, I. "Modeling slump of concrete with fly ash and superplasticizer." Computers and Concrete5.6 (2008): 559–572.
Gencel, Osman; et al. (2011). "Comparison of artificial neural networks and general linear model approaches for the analysis of abrasive wear of concrete". Construction and Building Materials. 25 (8): 3486–3494. doi:10.1016/j.conbuildmat.2011.03.040. /wiki/Doi_(identifier)
Dietterich, Thomas G., et al. "A comparison of dynamic reposing and tangent distance for drug activity prediction Archived 7 December 2019 at the Wayback Machine." Advances in Neural Information Processing Systems (1994): 216–216. http://papers.nips.cc/paper/781-a-comparison-of-dynamic-reposing-and-tangent-distance-for-drug-activity-prediction.pdf
Buscema, Massimo; Tastle, William J.; Terzi, Stefano (2013). "Meta Net: A New Meta-Classifier Family". Data Mining Applications Using Artificial Adaptive Systems. pp. 141–182. doi:10.1007/978-1-4614-4223-3_5. ISBN 978-1-4614-4222-6. 978-1-4614-4222-6
Barnard, Amanda; Sun, Baichuan; Motevalli Soumehsaraei, Ben; & Opletal, George (2019): Silver Nanoparticle Data Set. v3. CSIRO. Data Collection. https://doi.org/10.25919/5d22d20bc543e https://data.csiro.au/collection/csiro:23472
Barnard, Amanda; Sun, Baichuan; & Opletal, George (2019): Platinum Nanoparticle Data Set. v2. CSIRO. Data Collection. https://doi.org/10.25919/5d3958d9bf5f7 https://data.csiro.au/collection/csiro:36491
Barnard, Amanda; & Opletal, George (2019): Gold Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/5d395ef9a4291 https://data.csiro.au/collection/csiro:40669
Barnard, Amanda; & Opletal, George (2019): Ruthenium Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/5e30b8fa67484 https://data.csiro.au/collection/csiro:42601
Barnard, Amanda; & Opletal, George (2019): Copper Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/5e30ba386311f https://data.csiro.au/collection/csiro:42598
Barnard, Amanda; & Opletal, George (2023): Palladium Nanoparticle Data Set. v2. CSIRO. Data Collection. https://doi.org/10.25919/epxd-8p61 https://doi.org/10.25919/epxd-8p61
Ting, Jonathan; Barnard, Amanda; Opletal, George (2023): AuCo Nanoparticle Data Set. v2. CSIRO. Data Collection. https://doi.org/10.25919/7h3x-1343 https://doi.org/10.25919/7h3x-1343
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PtCo Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/jzh8-rd31 https://doi.org/10.25919/jzh8-rd31
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PtAu Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/tdnv-jp30 https://doi.org/10.25919/tdnv-jp30
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PdPt Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/qced-2e85 https://doi.org/10.25919/qced-2e85
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PdCo Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/az9t-vr97 https://doi.org/10.25919/az9t-vr97
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): CoPt Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/0bs4-sn79 https://doi.org/10.25919/0bs4-sn79
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): CoPd Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/em3a-9a89 https://doi.org/10.25919/em3a-9a89
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): CoAu Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/991j-hg07 https://doi.org/10.25919/991j-hg07
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): AuPt Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/7zh9-3f67 https://doi.org/10.25919/7zh9-3f67
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PtPd Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/9sz9-3a85 https://doi.org/10.25919/9sz9-3a85
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): PdAu Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/6ajg-1275 https://doi.org/10.25919/6ajg-1275
Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): AuPd Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/v0r5-sw08 https://doi.org/10.25919/v0r5-sw08
Lu, Kaihan; Ting, Jonathan; Barnard, Amanda; & Opletal, George (2023): AuPdPt Nanoparticle Data Set. v1. CSIRO. Data Collection. https://doi.org/10.25919/psvw-am47 https://doi.org/10.25919/psvw-am47
Amoradnejad, Issa; Amoradnejad, Rahimberdi; et al. (2022). "Age dataset: A structured general-purpose dataset on life, work, and death of 1.22 million distinguished people". Workshop Proceedings of the 16th International AAAI Conference on Web and Social Media (ICWSM). 3. ICWSM: 1–4. doi:10.36190/2022.82. S2CID 249668669. http://workshop-proceedings.icwsm.org/abstract?id=2022_82
"Age Dataset". GitHub. 7 June 2022. https://github.com/Moradnejad/AgeDataset
"Synthetic Fundus Dataset". Archived from the original on 29 November 2021. Retrieved 22 February 2023. https://web.archive.org/web/20211129155047/http://math.unipa.it/cvalenti/fundus/
Lo Castro, Dario; et al. (2020). "A visual framework to create photorealistic retinal vessels for diagnosis purposes". Journal of Biomedical Informatics. 108: 103490. doi:10.1016/j.jbi.2020.103490. PMID 32640292. S2CID 220429697. /wiki/Doi_(identifier)
Ingber, Lester (1997). "Statistical mechanics of neocortical interactions: Canonical momenta indicatorsof electroencephalography". Physical Review E. 55 (4): 4578–4593. arXiv:physics/0001052. Bibcode:1997PhRvE..55.4578I. doi:10.1103/PhysRevE.55.4578. S2CID 6390999. /wiki/ArXiv_(identifier)
Hoffmann, Ulrich; Vesin, Jean-Marc; Ebrahimi, Touradj; Diserens, Karin (January 2008). "An efficient P300-based brain–computer interface for disabled subjects". Journal of Neuroscience Methods. 167 (1): 115–125. doi:10.1016/j.jneumeth.2007.03.005. PMID 17445904. http://infoscience.epfl.ch/record/101093
Donchin, Emanuel; Spencer, Kevin M.; Wijesinghe, Ranjith (2000). "The mental prosthesis: assessing the speed of a P300-based brain-computer interface". IEEE Transactions on Rehabilitation Engineering. 8 (2): 174–179. doi:10.1109/86.847808. PMID 10896179. S2CID 84043. /wiki/Doi_(identifier)
Detrano, Robert; et al. (1989). "International application of a new probability algorithm for the diagnosis of coronary artery disease". The American Journal of Cardiology. 64 (5): 304–310. doi:10.1016/0002-9149(89)90524-9. PMID 2756873. /wiki/Doi_(identifier)
Bradley, Andrew P (1997). "The use of the area under the ROC curve in the evaluation of machine learning algorithms" (PDF). Pattern Recognition. 30 (7): 1145–1159. Bibcode:1997PatRe..30.1145B. doi:10.1016/s0031-3203(96)00142-2. S2CID 13806304. http://espace.library.uq.edu.au/view/UQ:8925/pr-t.pdf
Street, W. N.; Wolberg, W. H.; Mangasarian, O. L. (1993). "Nuclear feature extraction for breast tumor diagnosis". In Acharya, Raj S.; Goldgof, Dmitry B. (eds.). Biomedical Image Processing and Biomedical Visualization. Vol. 1905. pp. 861–870. doi:10.1117/12.148698. /wiki/Doi_(identifier)
Demir, Cigdem; Yener, Bülent (2005). Automated cancer diagnosis based on histopathological images : a systematic survey (PDF) (Report). S2CID 8952443. https://www.cs.rpi.edu/research/pdf/05-09.pdf
Abuse, Substance. "Mental Health Services Administration, Results from the 2010 National Survey on Drug Use and Health: Summary of National Findings, NSDUH Series H-41, HHS Publication No.(SMA) 11-4658." Rockville, MD: Substance Abuse and Mental Health Services Administration 201 (2011).
Hong, Zi-Quan; Yang, Jing-Yu (1991). "Optimal discriminant plane for a small number of samples and design method of classifier on the plane". Pattern Recognition. 24 (4): 317–324. Bibcode:1991PatRe..24..317H. doi:10.1016/0031-3203(91)90074-f. /wiki/Bibcode_(identifier)
Li, Jinyan; Wong, Limsoon (2003). "Using Rules to Analyse Bio-medical Data: A Comparison between C4.5 and PCL". Advances in Web-Age Information Management. Lecture Notes in Computer Science. Vol. 2762. pp. 254–265. doi:10.1007/978-3-540-45160-0_25. ISBN 978-3-540-40715-7. 978-3-540-40715-7
Guvenir, H.A.; Acar, B.; Demiroz, G.; Cekin, A. (1997). "A supervised machine learning algorithm for arrhythmia analysis". Computers in Cardiology 1997. pp. 433–436. doi:10.1109/CIC.1997.647926. hdl:11693/27699. ISBN 0-7803-4445-6. 0-7803-4445-6
Lagus, Krista; Alhoniemi, Esa; Seppä, Jeremias; Honkela, Antti; Wagner, Paul (2005). "Independent Variable Group Analysis in Learning Compact Representations for Data" (PDF). International and Interdisciplinary Conference on Adaptive Knowledge Representation and Reasoning (AKRR'05), Helsinki, Finland, June 15-17, 2005. pp. 49–56. http://research.ics.aalto.fi/events/AKRR05/papers/akrr05lagus.pdf
Strack, Beata, et al. "Impact of HbA1c measurement on hospital readmission rates: analysis of 70,000 clinical database patient records." BioMed Research International 2014; 2014 http://downloads.hindawi.com/journals/bmri/2014/781670.pdf
Rubin, Daniel J (2015). "Hospital readmission of patients with diabetes". Current Diabetes Reports. 15 (4): 1–9. doi:10.1007/s11892-015-0584-7. PMID 25712258. S2CID 3908599. /wiki/Doi_(identifier)
Antal, Bálint; Hajdu, András (2014). "An ensemble-based system for automatic screening of diabetic retinopathy". Knowledge-Based Systems. 60 (2014): 20–27. arXiv:1410.8576. Bibcode:2014arXiv1410.8576A. doi:10.1016/j.knosys.2013.12.023. S2CID 13984326. /wiki/ArXiv_(identifier)
Haloi, Mrinal (2015). "Improved Microaneurysm Detection using Deep Neural Networks". arXiv:1505.04424 [cs.CV]. /wiki/ArXiv_(identifier)
ELIE, Guillaume PATRY, Gervais GAUTHIER, Bruno LAY, Julien ROGER, Damien. "ADCIS Download Third Party: Messidor Database". adcis.net. Retrieved 25 February 2018.{{cite web}}: CS1 maint: multiple names: authors list (link) http://www.adcis.net/en/Download-Third-Party/Messidor.htmldownload.php
Decencière, Etienne; Zhang, Xiwei; Cazuguel, Guy; Lay, Bruno; Cochener, Béatrice; Trone, Caroline; Gain, Philippe; Ordonez, Richard; Massin, Pascale; Erginay, Ali; Charton, Béatrice; Klein, Jean-Claude (26 August 2014). "Feedback on a Publicly Distributed Image Database: The Messidor Database". Image Analysis & Stereology. 33 (3): 231. doi:10.5566/ias.1155. /wiki/Doi_(identifier)
Bagirov, A. M.; Rubinov, A. M.; Soukhoroukova, N. V.; Yearwood, J. (June 2003). "Unsupervised and supervised data classification via nonsmooth and global optimization". Top. 11 (1): 1–75. doi:10.1007/bf02578945. https://figshare.com/articles/journal_contribution/26292709
Fung, Glenn; Dundar, Murat; Bi, Jinbo; Rao, Bharat (2004). "A fast iterative algorithm for fisher discriminant using heterogeneous kernels". In Greiner, Russell; Schuurmans, Dale (eds.). Proceedings of the Twenty-first International Conference on Machine Learning. ACM. p. 40. doi:10.1145/1015330.1015409. ISBN 978-1-58113-838-2. 978-1-58113-838-2
Quinlan, J. R.; Compton, P. J.; Horn, K. A.; Lazarus, L. (1987). "Inductive knowledge acquisition: a case study". In Quinlan, John Ross (ed.). Applications of Expert Systems: Based on the Proceedings of the Second Australian Conference. Turing Institute Press. pp. 137–156. ISBN 978-0-201-17449-6. 978-0-201-17449-6
Zhi-Hua Zhou; Yuan Jiang (2004). "NeC4.5: Neural ensemble based C4.5". IEEE Transactions on Knowledge and Data Engineering. 16 (6): 770–773. doi:10.1109/tkde.2004.11. /wiki/Doi_(identifier)
Er, Orhan; et al. (2012). "An approach based on probabilistic neural network for diagnosis of Mesothelioma's disease". Computers & Electrical Engineering. 38 (1): 75–81. doi:10.1016/j.compeleceng.2011.09.001. /wiki/Doi_(identifier)
Er, Orhan; Tanrikulu, A. Çetin; Abakay, Abdurrahman (10 May 2015). "Use of artificial intelligence techniques for diagnosis of malignant pleural mesothelioma". Dicle Medical Journal / Dicle Tip Dergisi. 42 (1). doi:10.5798/diclemedj.0921.2015.01.0520 (inactive 23 November 2024).{{cite journal}}: CS1 maint: DOI inactive as of November 2024 (link) /wiki/Doi_(identifier)
Li, Michael H.; Mestre, Tiago A.; Fox, Susan H.; Taati, Babak (25 July 2017). "Vision-Based Assessment of Parkinsonism and Levodopa-Induced Dyskinesia with Deep Learning Pose Estimation". Journal of Neuroengineering and Rehabilitation. 15 (1): 97. arXiv:1707.09416. Bibcode:2017arXiv170709416L. doi:10.1186/s12984-018-0446-z. PMC 6219082. PMID 30400914. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6219082
Li, Michael H.; Mestre, Tiago A.; Fox, Susan H.; Taati, Babak (August 2018). "Automated assessment of levodopa-induced dyskinesia: Evaluating the responsiveness of video-based features". Parkinsonism & Related Disorders. 53: 42–45. doi:10.1016/j.parkreldis.2018.04.036. PMID 29748112. /wiki/Doi_(identifier)
"Parkinson's Vision-Based Pose Estimation Dataset | Kaggle". kaggle.com. Retrieved 22 August 2018. https://www.kaggle.com/limi44/parkinsons-visionbased-pose-estimation-dataset/home
Shannon, Paul; et al. (2003). "Cytoscape: a software environment for integrated models of biomolecular interaction networks". Genome Research. 13 (11): 2498–2504. doi:10.1101/gr.1239303. PMC 403769. PMID 14597658. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC403769
Javadi, Soroush; Mirroshandel, Seyed Abolghasem (June 2019). "A novel deep learning method for automatic assessment of human sperm images". Computers in Biology and Medicine. 109: 182–194. doi:10.1016/j.compbiomed.2019.04.030. PMID 31059902. /wiki/Doi_(identifier)
"soroushj/mhsma-dataset: MHSMA: The Modified Human Sperm Morphology Analysis Dataset". github.com. Retrieved 3 May 2019. https://github.com/soroushj/mhsma-dataset
Clark, David, Zoltan Schreter, and Anthony Adams. "A quantitative comparison of dystal and backpropagation." Proceedings of 1996 Australian Conference on Neural Networks. 1996.
Jiang, Yuan, and Zhi-Hua Zhou. "Editing training data for kNN classifiers with neural network ensemble." Advances in Neural Networks–ISNN 2004. Springer Berlin Heidelberg, 2004. 356–361. https://cs.nju.edu.cn/zhouzh/zhouzh.files/publication/isnn04a.pdf
Ontañón, Santiago; Plaza, Enric (2009). "On Similarity Measures Based on a Refinement Lattice". Case-Based Reasoning Research and Development. Lecture Notes in Computer Science. Vol. 5650. pp. 240–255. doi:10.1007/978-3-642-02998-1_18. ISBN 978-3-642-02997-4. 978-3-642-02997-4
"PLF data inventory". GitHub. 5 November 2021. https://github.com/Animal-Data-Inventory/PLFDataInventory
Li, Jinyan; Wong, Limsoon (2003). "Using Rules to Analyse Bio-medical Data: A Comparison between C4.5 and PCL". Advances in Web-Age Information Management. Lecture Notes in Computer Science. Vol. 2762. pp. 254–265. doi:10.1007/978-3-540-45160-0_25. ISBN 978-3-540-40715-7. 978-3-540-40715-7
Higuera, Clara; Gardiner, Katheleen J.; Cios, Krzysztof J. (2015). "Self-organizing feature maps identify proteins critical to learning in a mouse model of down syndrome". PLOS ONE. 10 (6): e0129126. Bibcode:2015PLoSO..1029126H. doi:10.1371/journal.pone.0129126. PMC 4482027. PMID 26111164. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4482027
Ahmed, Md Mahiuddin; et al. (2015). "Protein dynamics associated with failed and rescued learning in the Ts65Dn mouse model of Down syndrome". PLOS ONE. 10 (3): e0119491. Bibcode:2015PLoSO..1019491A. doi:10.1371/journal.pone.0119491. PMC 4368539. PMID 25793384. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4368539
Langley, PAT (2014). "Trading off simplicity and coverage in incremental concept learning" (PDF). Machine Learning Proceedings. 1988: 73. Archived from the original (PDF) on 6 August 2019. Retrieved 6 August 2019. https://web.archive.org/web/20190806184005/https://www.westmont.edu/~iba/pubs/hillary-paper.pdf
"Mushroom Data Set 2020". mushroom.mathematik.uni-marburg.de. Retrieved 6 April 2021. https://mushroom.mathematik.uni-marburg.de/
Wagner, Dennis; Heider, Dominik; Hattab, Georges (14 April 2021). "Mushroom data creation, curation, and simulation to support classification tasks". Scientific Reports. 11 (1): 8134. Bibcode:2021NatSR..11.8134W. doi:10.1038/s41598-021-87602-3. PMC 8046754. PMID 33854157. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8046754
Cortez, Paulo, and Aníbal de Jesus Raimundo Morais. "A data mining approach to predict forest fires using meteorological data." (2007).
Farquad, M. A. H.; Ravi, V.; Raju, S. Bapi (2010). "Support vector regression based hybrid rule extraction methods for forecasting". Expert Systems with Applications. 37 (8): 5577–5589. doi:10.1016/j.eswa.2010.02.055. /wiki/Doi_(identifier)
Fisher, Ronald A (1936). "The use of multiple measurements in taxonomic problems". Annals of Eugenics. 7 (2): 179–188. doi:10.1111/j.1469-1809.1936.tb02137.x. hdl:2440/15227. /wiki/Doi_(identifier)
Ghahramani, Zoubin, and Michael I. Jordan. "Supervised learning from incomplete data via an EM approach Archived 22 April 2017 at the Wayback Machine." Advances in neural information processing systems 6. 1994. http://papers.nips.cc/paper/767-supervised-learning-from-incomplete-data-via-an-em-approach.pdf
Mallah, Charles; Cope, James; Orwell, James (2013). "Plant Leaf Classification using Probabilistic Integration of Shape, Texture and Margin Features". Computer Graphics and Imaging / 798: Signal Processing, Pattern Recognition and Applications. doi:10.2316/P.2013.798-098. ISBN 978-0-88986-944-8. 978-0-88986-944-8
Yahiaoui, Itheri; Mzoughi, Olfa; Boujemaa, Nozha (2012). "Leaf Shape Descriptor for Tree Species Identification". 2012 IEEE International Conference on Multimedia and Expo. pp. 254–259. doi:10.1109/ICME.2012.130. ISBN 978-1-4673-1659-0. 978-1-4673-1659-0
Tan, Ming; Eshelman, Larry (1988). "Using Weighted Networks to Represent Classification Knowledge in Noisy Domains". Machine Learning Proceedings 1988. pp. 121–134. doi:10.1016/B978-0-934613-64-4.50018-9. ISBN 978-0-934613-64-4. 978-0-934613-64-4
Charytanowicz, Małgorzata, et al. "Complete gradient clustering algorithm for features analysis of x-ray images." Information technologies in biomedicine. Springer Berlin Heidelberg, 2010. 15–24. http://home.agh.edu.pl/~kulpi/publ/Charytanowicz_Niewczas_Kulczycki_Kowalski_Lukasik_Zak_-_Information_Technologies_in_Biomedicine_-_2010.pdf
Sanchez, Mauricio A.; et al. (2014). "Fuzzy granular gravitational clustering algorithm for multivariate data". Information Sciences. 279: 498–511. doi:10.1016/j.ins.2014.04.005. /wiki/Doi_(identifier)
Blackard, Jock A.; Dean, Denis J. (December 1999). "Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables". Computers and Electronics in Agriculture. 24 (3): 131–151. Bibcode:1999CEAgr..24..131B. doi:10.1016/s0168-1699(99)00046-0. /wiki/Bibcode_(identifier)
Fürnkranz, Johannes (2001). "Round Robin Rule Learning" (PDF). In Danyluk, Andrea Pohoreckyj; Brodley, Carla E. (eds.). Machine Learning: Proceedings of the Eighteenth International Conference (ICML 2001) : Williams College, June 28-July 1, 2001. Morgan Kaufmann Publishers. pp. 146–153. ISBN 978-1-55860-778-1. 978-1-55860-778-1
Li, Song; Assmann, Sarah M.; Albert, Réka (2006). "Predicting essential components of signal transduction networks: a dynamic model of guard cell abscisic acid signaling". PLOS Biol. 4 (10): e312. arXiv:q-bio/0610012. Bibcode:2006q.bio....10012L. doi:10.1371/journal.pbio.0040312. PMC 1564158. PMID 16968132. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC1564158
Munisami, Trishen; et al. (2015). "Plant Leaf Recognition Using Shape Features and Colour Histogram with K-nearest Neighbour Classifiers". Procedia Computer Science. 58: 740–747. doi:10.1016/j.procs.2015.08.095. https://doi.org/10.1016%2Fj.procs.2015.08.095
Li, Bai (2016). "Atomic potential matching: An evolutionary target recognition approach based on edge features". Optik. 127 (5): 3162–3168. Bibcode:2016Optik.127.3162L. doi:10.1016/j.ijleo.2015.11.186. /wiki/Bibcode_(identifier)
Razavian, Ali, et al. "CNN features off-the-shelf: an astounding baseline for recognition." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2014. https://www.cv-foundation.org/openaccess/content_cvpr_workshops_2014/W15/papers/Razavian_CNN_Features_Off-the-Shelf_2014_CVPR_paper.pdf
Nilsback, Maria-Elena, and Andrew Zisserman. "A visual vocabulary for flower classification."Computer Vision and Pattern Recognition, 2006 IEEE Computer Society Conference on. Vol. 2. IEEE, 2006. http://www.robots.ox.ac.uk/~men/papers/nilsback_cvpr06.pdf
Giselsson, Thomas M.; et al. (2017). "A Public Image Database for Benchmark of Plant Seedling Classification Algorithms". arXiv:1711.05458 [cs.CV]. /wiki/ArXiv_(identifier)
Oltean, Mihai (2017). "Fruits-360 dataset". GitHub. https://www.github.com/fruits-360
Old, Richard (2024). "Weed-ID.App dataset". https://weed-id.app
Rahman, Abdur; Lu, Yuzhen; Wang, Haifeng (February 2023). "Performance evaluation of deep learning object detectors for weed detection for cotton". Smart Agricultural Technology. 3: 100126. doi:10.1016/j.atech.2022.100126. /wiki/Doi_(identifier)
Nakai, Kenta; Kanehisa, Minoru (1991). "Expert system for predicting protein localization sites in gram-negative bacteria". Proteins: Structure, Function, and Bioinformatics. 11 (2): 95–110. doi:10.1002/prot.340110203. PMID 1946347. S2CID 27606447. /wiki/Doi_(identifier)
Ling, Charles X., et al. "Decision trees with minimal costs." Proceedings of the twenty-first international conference on Machine learning. ACM, 2004. https://cling.csd.uwo.ca/cs860/ICML04-Ling.pdf
Mahé, Pierre; Arsac, Maud; Chatellier, Sonia; Monnin, Valérie; Perrot, Nadine; Mailler, Sandrine; Girard, Victoria; Ramjeet, Mahendrasingh; Surre, Jérémy; Lacroix, Bruno; van Belkum, Alex; Veyrieras, Jean-Baptiste (May 2014). "Automatic identification of mixed bacterial species fingerprints in a MALDI-TOF mass-spectrum". Bioinformatics. 30 (9): 1280–1286. doi:10.1093/bioinformatics/btu022. PMID 24443381. /wiki/Doi_(identifier)
Barbano, Duane; et al. (2015). "Rapid characterization of microalgae and microalgae mixtures using matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS)". PLOS ONE. 10 (8): e0135337. Bibcode:2015PLoSO..1035337B. doi:10.1371/journal.pone.0135337. PMC 4536233. PMID 26271045. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4536233
Horton, Paul; Nakai, Kenta (1996). "A probabilistic classification system for predicting the cellular localization sites of proteins" (PDF). ISMB-96 Proceedings. 4: 109–15. PMID 8877510. Archived from the original (PDF) on 4 November 2021. Retrieved 6 August 2019. https://web.archive.org/web/20211104042943/https://www.aaai.org/Papers/ISMB/1996/ISMB96-012.pdf
Allwein, Erin L.; Schapire, Robert E.; Singer, Yoram (2001). "Reducing multiclass to binary: A unifying approach for margin classifiers" (PDF). The Journal of Machine Learning Research. 1: 113–141. http://www.jmlr.org/papers/volume1/allwein00a/allwein00a.pdf
Mayr, Andreas; Klambauer, Guenter; Unterthiner, Thomas; Hochreiter, Sepp (2016). "DeepTox: Toxicity Prediction Using Deep Learning". Frontiers in Environmental Science. 3: 80. Bibcode:2016FrEnS...3...80M. doi:10.3389/fenvs.2015.00080. http://bioinf.jku.at/research/DeepTox/tox21.html
Lavin, Alexander; Ahmad, Subutai (12 October 2015). "Evaluating Real-Time Anomaly Detection Algorithms -- the Numenta Anomaly Benchmark". 2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA). pp. 38–44. arXiv:1510.03336. doi:10.1109/ICMLA.2015.141. ISBN 978-1-5090-0287-0. S2CID 6842305. 978-1-5090-0287-0
Iurii D. Katser; Vyacheslav O. Kozitsin. "SKAB GitHub repository". GitHub. Retrieved 12 January 2021. https://github.com/waico/skab
Iurii D. Katser; Vyacheslav O. Kozitsin (2020). "Skoltech Anomaly Benchmark (SKAB)". Kaggle. doi:10.34740/KAGGLE/DSV/1693952. Retrieved 12 January 2021. {{cite journal}}: Cite journal requires |journal= (help) https://www.kaggle.com/yuriykatser/skoltech-anomaly-benchmark-skab
Campos, Guilherme O.; Zimek, Arthur; Sander, Jörg; Campello, Ricardo J. G. B.; Micenková, Barbora; Schubert, Erich; Assent, Ira; Houle, Michael E. (July 2016). "On the evaluation of unsupervised outlier detection: measures, datasets, and an empirical study". Data Mining and Knowledge Discovery. 30 (4): 891–927. doi:10.1007/s10618-015-0444-8. /wiki/Doi_(identifier)
Ann-Kathrin Hartmann, Tommaso Soru, Edgard Marx. Generating a Large Dataset for Neural Question Answering over the DBpedia Knowledge Base. 2018. https://www.researchgate.net/publication/324482598_Generating_a_Large_Dataset_for_Neural_Question_Answering_over_the_DBpedia_Knowledge_Base
Soru, Tommaso; Marx, Edgard; Moussallem, Diego; Publio, Gustavo; Valdestilhas, André; Esteves, Diego; Neto, Ciro Baron (2017). SPARQL as a Foreign Language (Preprint). arXiv:1708.07624. /wiki/ArXiv_(identifier)
Kiet Van Nguyen, Duc-Vu Nguyen, Anh Gia-Tuan Nguyen, Ngan Luu-Thuy Nguyen. A Vietnamese Dataset for Evaluating Machine Reading Comprehension. COLING 2020. https://www.aclweb.org/anthology/2020.coling-main.233.pdf
Nguyen, Kiet Van; Tran, Khiem Vinh; Luu, Son T.; Nguyen, Anh Gia-Tuan; Nguyen, Ngan Luu-Thuy (2020). "Enhancing Lexical-Based Approach With External Knowledge for Vietnamese Multiple-Choice Machine Reading Comprehension". IEEE Access. 8: 201404–201417. Bibcode:2020IEEEA...8t1404N. doi:10.1109/ACCESS.2020.3035701. /wiki/Bibcode_(identifier)
Anantha, Raviteja; Vakulenko, Svitlana; Tu, Zhucheng; Longpre, Shayne; Pulman, Stephen; Chappidi, Srinivas (2020). "Open-Domain Question Answering Goes Conversational via Question Rewriting". arXiv:2010.04898 [cs.IR]. /wiki/ArXiv_(identifier)
Khashabi, Daniel; Min, Sewon; Khot, Tushar; Sabharwal, Ashish; Tafjord, Oyvind; Clark, Peter; Hajishirzi, Hannaneh (November 2020). "UNIFIEDQA: Crossing Format Boundaries with a Single QA System". Findings of the Association for Computational Linguistics: EMNLP 2020. Online: Association for Computational Linguistics: 1896–1907. arXiv:2005.00700. doi:10.18653/v1/2020.findings-emnlp.171. S2CID 218487109. https://aclanthology.org/2020.findings-emnlp.171
Taskmaster, Google Research Datasets, 17 December 2022, retrieved 7 January 2023 https://github.com/google-research-datasets/Taskmaster
Byrne, Bill; Krishnamoorthi, Karthik; Sankar, Chinnadhurai; Neelakantan, Arvind; Duckworth, Daniel; Yavuz, Semih; Goodrich, Ben; Dubey, Amit; Cedilnik, Andy; Kim, Kyu-Young (1 September 2019). "Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset". arXiv:1909.05358 [cs.CL]. /wiki/ArXiv_(identifier)
Yasunaga, Michihiro; Liang, Percy (21 November 2020). "Graph-based, Self-Supervised Program Repair from Diagnostic Feedback". International Conference on Machine Learning. PMLR: 10799–10808. arXiv:2005.10636. https://proceedings.mlr.press/v119/yasunaga20a.html
Wang, Yizhong; Mishra, Swaroop; Alipoormolabashi, Pegah; Kordi, Yeganeh; Mirzaei, Amirreza; Arunkumar, Anjana; Ashok, Arjun; Dhanasekaran, Arut Selvan; Naik, Atharva; Stap, David; Pathak, Eshaan; Karamanolakis, Giannis; Lai, Haizhi Gary; Purohit, Ishan; Mondal, Ishani (24 October 2022). "Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks". arXiv:2204.07705 [cs.CL]. /wiki/ArXiv_(identifier)
Paperno, Denis; Kruszewski, Germán; Lazaridou, Angeliki; Pham, Quan Ngoc; Bernardi, Raffaella; Pezzelle, Sandro; Baroni, Marco; Boleda, Gemma; Fernández, Raquel (7 August 2016), The LAMBADA dataset, doi:10.5281/zenodo.2630551, retrieved 7 January 2023 https://zenodo.org/record/2630551
Paperno, Denis; Kruszewski, Germán; Lazaridou, Angeliki; Pham, Ngoc Quan; Bernardi, Raffaella; Pezzelle, Sandro; Baroni, Marco; Boleda, Gemma; Fernández, Raquel (August 2016). "The LAMBADA dataset: Word prediction requiring a broad discourse context". Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Berlin, Germany: Association for Computational Linguistics: 1525–1534. doi:10.18653/v1/P16-1144. hdl:10230/32702. S2CID 2381275. https://aclanthology.org/P16-1144
Wei, Jason; Bosma, Maarten; Zhao, Vincent; Guu, Kelvin; Yu, Adams Wei; Lester, Brian; Du, Nan; Dai, Andrew M.; Le, Quoc V. (10 February 2022). Finetuned Language Models are Zero-Shot Learners (Preprint). arXiv:2109.01652. https://openreview.net/forum?id=gEZrGCozdqR
"Working with ATT&CK | MITRE ATT&CK®". attack.mitre.org. Retrieved 14 January 2023. https://attack.mitre.org/resources/working-with-attack/
"CAPEC - Common Attack Pattern Enumeration and Classification (CAPEC™)". capec.mitre.org. Retrieved 14 January 2023. https://capec.mitre.org/
"CVE - Home". cve.mitre.org. Retrieved 14 January 2023. https://cve.mitre.org/cve/
"CWE - Common Weakness Enumeration". cwe.mitre.org. Retrieved 14 January 2023. https://cwe.mitre.org/index.html
Lim, Swee Kiat; Muis, Aldrian Obaja; Lu, Wei; Ong, Chen Hui (July 2017). "MalwareTextDB: A Database for Annotated Malware Articles". Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vancouver, Canada: Association for Computational Linguistics: 1557–1567. doi:10.18653/v1/P17-1143. S2CID 7816596. https://aclanthology.org/P17-1143
"USENIX". USENIX. Retrieved 19 January 2023. https://www.usenix.org/
"APTnotes | Read the Docs". readthedocs.org. Retrieved 19 January 2023. https://readthedocs.org/projects/aptnotes/
"Cryptography and Security authors/titles recent submissions". arxiv.org. Retrieved 19 January 2023. https://arxiv.org/list/cs.CR/recent
"Holistic Info-Sec for Web Developers - Fascicle 0". f0.holisticinfosecforwebdevelopers.com. Retrieved 20 January 2023. https://f0.holisticinfosecforwebdevelopers.com/
"Holistic Info-Sec for Web Developers - Fascicle 1". f1.holisticinfosecforwebdevelopers.com. Retrieved 20 January 2023. https://f1.holisticinfosecforwebdevelopers.com/
Vincent, Adam. "Web Services Web Services Hacking and Hardening" (PDF). owasp.org. https://owasp.org/www-pdf-archive/Web_Services_Hacking_and_Hardening.pdf
McCray, Joe. "Advanced SQL Injection" (PDF). defcon.org. https://defcon.org/images/defcon-17/dc-17-presentations/defcon-17-joseph_mccray-adv_sql_injection.pdf
Shah, Shreeraj. "Blind SQL injection discovery & exploitation technique" (PDF). blueinfy.com. https://blueinfy.com/wp/blindsql.pdf
Palcer, C. C. "Ethical hacking" (PDF). textfiles. https://blueinfy.com/wp/blindsql.pdf
"Hacking Secrets Revealed - Information and Instructional Guide" (PDF). https://www.onlinepot.org/security/HackersSecrets.pdf
Park, Alexis. "Hack any website" (PDF). https://defcon.org/images/defcon-11/dc-11-presentations/dc-11-Gentil/dc-11-gentil.pdf
Cerrudo, Cesar; Martinez Fayo, Esteban. "Hacking Databases for Owning your Data" (PDF). blackhat. https://www.blackhat.com/presentations/bh-europe-07/Cerrudo/Whitepaper/bh-eu-07-cerrudo-WP-up.pdf
O'Connor, Tj. "Violent Python-A Cookbook for Hackers, Forensic Analysts, Penetration Testers and Security Engineers" (PDF). Github. https://github.com/reconSF/python/blob/master/Syngress.Violent.Python.a.Cookbook.for.Hackers.2013.pdf
Grand, Joe. "Hardware Reverse Engineering: Access, Analyze, & Defeat" (PDF). blackhat. https://media.blackhat.com/bh-dc-11/Grand/BlackHat_DC_2011_Grand-Workshop.pdf
Chang, Jason V. "Computer Hacking: Making the Case for National Reporting Requirement" (PDF). cyber.harvard.edu. https://cyber.harvard.edu/sites/cyber.law.harvard.edu/files/ComputerHacking.pdf
"National Cybersecurity Strategies Repository". ITU. Retrieved 20 January 2023. https://www.itu.int:443/en/ITU-D/Cybersecurity/Pages/National-Strategies-repository.aspx
Chen, Yanlin (31 August 2022), Cyber Security Natural Language Processing, retrieved 20 January 2023 https://github.com/Ychen463/Cyber
Zampieri, Marcos; Malmasi, Shervin; Nakov, Preslav; Rosenthal, Sara; Farra, Noura; Kumar, Ritesh (16 April 2019). "Predicting the Type and Target of Offensive Posts in Social Media". arXiv:1902.09666 [cs.CL]. /wiki/ArXiv_(identifier)
"Threat reports". www.ncsc.gov.uk. Retrieved 20 January 2023. https://www.ncsc.gov.uk/section/keep-up-to-date/threat-reports
"Category: APT reports | Securelist". securelist.com. Retrieved 23 January 2023. https://securelist.com/category/apt-reports/
"Your Cybersecurity News Connection - Cyber News | CyberWire". The CyberWire. Retrieved 23 January 2023. https://thecyberwire.com/
"News". 21 August 2016. Retrieved 23 January 2023. https://www.databreaches.net/news/
"Cybernews". Cybernews. https://cybernews.com/
"BleepingComputer". BleepingComputer. Retrieved 23 January 2023. https://www.bleepingcomputer.com/
"Homepage". The Record from Recorded Future News. Retrieved 23 January 2023. https://therecord.media/
"HackRead | Latest Cyber Crime - InfoSec- Tech - Hacking News". 8 January 2022. Retrieved 23 January 2023. https://www.hackread.com/
"Securelist | Kaspersky's threat research and reports". securelist.com. Retrieved 31 January 2023. https://securelist.com/
Harshaw, Christopher R.; Bridges, Robert A.; Iannacone, Michael D.; Reed, Joel W.; Goodall, John R. (5 April 2016). "GraphPrints". Proceedings of the 11th Annual Cyber and Information Security Research Conference. CISRC '16. New York, NY, USA: Association for Computing Machinery. pp. 1–4. doi:10.1145/2897795.2897806. ISBN 978-1-4503-3752-6. 978-1-4503-3752-6
"Farsight Security, cyber security intelligence solutions". Farsight Security. Retrieved 13 February 2023. https://www.farsightsecurity.com/
"Schneier on Security". www.schneier.com. Retrieved 13 February 2023. https://www.schneier.com/
"#1 in Cloud Security & Endpoint Cybersecurity". Trend Micro. Retrieved 13 February 2023. https://www.trendmicro.com/en_us/business.html
"The Hacker News | #1 Trusted Cybersecurity News Site". The Hacker News. Retrieved 13 February 2023. https://thehackernews.com/
"Krebs on Security – In-depth security news and investigation". Retrieved 25 February 2023. https://krebsonsecurity.com/
"MITRE D3FEND Knowledge Graph". d3fend.mitre.org. Retrieved 31 March 2023. https://d3fend.mitre.org/
"MITRE | ATLAS™". atlas.mitre.org. Retrieved 31 March 2023. https://atlas.mitre.org/
"MITRE Engage™ | An Adversary Engagement Framework from MITRE". Retrieved 1 April 2023. https://engage.mitre.org/
"Hacking Tutorials - The best Step-by-Step Hacking Tutorials". Hacking Tutorials. Retrieved 1 April 2023. https://www.hackingtutorials.org/
"TCFD Knowledge Hub". TCFD Knowledge Hub. Retrieved 3 February 2023. https://www.tcfdhub.org/
"ResponsibilityReports.com". www.responsibilityreports.com. Retrieved 3 February 2023. https://www.responsibilityreports.com/
"About — IPCC". Retrieved 20 February 2023. https://www.ipcc.ch/about/
"Alliance for Research on Corporate Sustainability | ARCS serves as a vehicle for advancing rigorous academic research on corporate sustainability issues". corporate-sustainability.org. Retrieved 2 March 2023. https://corporate-sustainability.org/
Mehra, Srishti; Louka, Robert; Zhang, Yixun (2022). "ESGBERT: Language Model to Help with Classification Tasks Related to Companies' Environmental, Social, and Governance Practices". Embedded Systems and Applications. pp. 183–190. doi:10.5121/csit.2022.120616. ISBN 978-1-925953-65-7. 978-1-925953-65-7
This article incorporates text available under the CC BY 4.0 license. //creativecommons.org/licenses/by/4.0/
Diggelmann, Thomas; Boyd-Graber, Jordan; Bulian, Jannis; Ciaramita, Massimiliano; Leippold, Markus (2 January 2021). "CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims". arXiv:2012.00614 [cs.CL]. /wiki/ArXiv_(identifier)
"climate-news-db". www.climate-news-db.com. Retrieved 3 February 2023. http://www.climate-news-db.com/
"Climatext". www.sustainablefinance.uzh.ch. Retrieved 19 February 2023. http://www.sustainablefinance.uzh.ch/en/research/climate-fever/climatext.html
"Greenbiz". www.greenbiz.com. Retrieved 2 March 2023. https://www.greenbiz.com/
"Explore the @Reuters Hot List of 1,000 top climate scientists". Reuters. Retrieved 22 March 2023. https://www.reuters.com/investigates/special-report/climate-change-scientists-list/
"Blogs | Alliance for Research on Corporate Sustainability". corporate-sustainability.org. Retrieved 27 March 2023. https://corporate-sustainability.org/blogs/
"Greenbiz". www.greenbiz.com. Retrieved 29 March 2023. https://www.greenbiz.com/
"CSR News". www.csrwire.com. Retrieved 29 March 2023. https://www.csrwire.com/press_releases
"CDP Homepage". www.cdp.net. Retrieved 29 March 2023. https://www.cdp.net/en
de Vries, Harm (2022). "The Stack: 3 TB of permissively licensed source code". arXiv:2211.15533 [cs.CL]. /wiki/ArXiv_(identifier)
"The Stack Dedup". Huggingface. Retrieved 29 August 2023. https://huggingface.co/datasets/bigcode/the-stack-dedup
"Hybrid cloud blog". content.cloud.redhat.com. Retrieved 9 April 2023. https://content.cloud.redhat.com/blog
"Production-Grade Container Orchestration". Kubernetes. Retrieved 9 April 2023. https://kubernetes.io/
"Home | Official Red Hat OpenShift Documentation". docs.openshift.com. Retrieved 9 April 2023. https://docs.openshift.com/
"Cloud Native Computing Foundation". Cloud Native Computing Foundation. Retrieved 9 April 2023. https://www.cncf.io/
CNCF Community Presentations, Cloud Native Computing Foundation (CNCF), 11 April 2023, retrieved 11 April 2023 https://github.com/cncf/presentations/blob/2ff57e4d78f6d70bb1fd5daf81e76f04a54c8520/kubernetes/README.md
"Red Hat - We make open source technologies for the enterprise". www.redhat.com. Retrieved 1 May 2023. https://www.redhat.com/en
Brown, Michael Scott; Pelosi, Michael J.; Dirska, Henry (2013). "Dynamic-Radius Species-Conserving Genetic Algorithm for the Financial Forecasting of Dow Jones Index Stocks". Machine Learning and Data Mining in Pattern Recognition. Lecture Notes in Computer Science. Vol. 7988. pp. 27–41. doi:10.1007/978-3-642-39712-7_3. ISBN 978-3-642-39711-0. 978-3-642-39711-0
Shen, Kao-Yi; Tzeng, Gwo-Hshiung (2015). "Fuzzy Inference-Enhanced VC-DRSA Model for Technical Analysis: Investment Decision Aid". International Journal of Fuzzy Systems. 17 (3): 375–389. doi:10.1007/s40815-015-0058-8. S2CID 68241024. /wiki/Doi_(identifier)
Quinlan, J.R. (September 1987). "Simplifying decision trees". International Journal of Man-Machine Studies. 27 (3): 221–234. doi:10.1016/s0020-7373(87)80053-6. hdl:1721.1/6453. /wiki/Doi_(identifier)
Hamers, Bart; Suykens, Johan AK; De Moor, Bart (2003). "Coupled transductive ensemble learning of kernel models" (PDF). Journal of Machine Learning Research. 1: 1–48. http://ftp.esat.kuleuven.be/pub/SISTA/hamers/BH_clm.pdf
Shmueli, Galit; Russo, Ralph P.; Jank, Wolfgang (December 2007). "The BARISTA: A model for bid arrivals in online auctions". The Annals of Applied Statistics. 1 (2). doi:10.1214/07-AOAS117. /wiki/Doi_(identifier)
Peng, Jie; Müller, Hans-Georg (September 2008). "Distance-based clustering of sparsely observed stochastic processes, with applications to online auctions". The Annals of Applied Statistics. 2 (3). doi:10.1214/08-AOAS172. /wiki/Doi_(identifier)
Eggermont, Jeroen; Kok, Joost N.; Kosters, Walter A. (2004). "Genetic Programming for data classification: Partitioning the search space". Proceedings of the 2004 ACM symposium on Applied computing. pp. 1001–1005. doi:10.1145/967900.968104. ISBN 978-1-58113-812-2. 978-1-58113-812-2
Moro, Sérgio; Cortez, Paulo; Rita, Paulo (2014). "A data-driven approach to predict the success of bank telemarketing". Decision Support Systems. 62: 22–31. doi:10.1016/j.dss.2014.03.001. hdl:10071/9499. S2CID 14181100. /wiki/Doi_(identifier)
Payne, Richard D.; Mallick, Bani K. (2014). "Bayesian Big Data Classification: A Review with Complements". arXiv:1411.5653 [stat.ME]. /wiki/ArXiv_(identifier)
Akbilgic, Oguz; Bozdogan, Hamparsum; Balaban, M. Erdal (2014). "A novel Hybrid RBF Neural Networks model as a forecaster". Statistics and Computing. 24 (3): 365–375. doi:10.1007/s11222-013-9375-7. S2CID 17764829. /wiki/Doi_(identifier)
Jabin, Suraiya (20 August 2014). "Stock Market Prediction using Feed-forward Artificial Neural Network". International Journal of Computer Applications. 99 (9): 4–8. Bibcode:2014IJCA...99i...4J. doi:10.5120/17399-7959. /wiki/Bibcode_(identifier)
Yeh, I-Cheng; Che-hui, Lien (2009). "The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients". Expert Systems with Applications. 36 (2): 2473–2480. doi:10.1016/j.eswa.2007.12.020. S2CID 15696161. /wiki/Doi_(identifier)
Lin, Shu Ling (2009). "A new two-stage hybrid approach of credit risk in banking industry". Expert Systems with Applications. 36 (4): 8333–8341. doi:10.1016/j.eswa.2008.10.015. /wiki/Doi_(identifier)
Xu, Yumo; Cohen, Shay B. (2018). "Stock Movement Prediction from Tweets and Historical Prices". Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). pp. 1970–1979. doi:10.18653/v1/P18-1183. /wiki/Doi_(identifier)
Pelckmans, Kristiaan; et al. (2005). "The differogram: Non-parametric noise variance estimation and its use for model selection". Neurocomputing. 69 (1): 100–122. doi:10.1016/j.neucom.2005.02.015. /wiki/Doi_(identifier)
Bay, Stephen D.; Kibler, Dennis; Pazzani, Michael J.; Smyth, Padhraic (December 2000). "The UCI KDD archive of large data sets for data mining research and experimentation". ACM SIGKDD Explorations Newsletter. 2 (2): 81–85. doi:10.1145/380995.381030. /wiki/Doi_(identifier)
Lucas, D. D.; et al. (2015). "Designing optimal greenhouse gas observing networks that consider performance and cost". Geoscientific Instrumentation, Methods and Data Systems. 4 (1): 121. Bibcode:2015GI......4..121L. doi:10.5194/gi-4-121-2015. https://doi.org/10.5194%2Fgi-4-121-2015
Pales, Jack C.; Keeling, Charles D. (1965). "The concentration of atmospheric carbon dioxide in Hawaii". Journal of Geophysical Research. 70 (24): 6053–6076. Bibcode:1965JGR....70.6053P. doi:10.1029/jz070i024p06053. /wiki/Bibcode_(identifier)
Zhi-Hua Zhou; Yuan Jiang (2004). "NeC4.5: Neural ensemble based C4.5". IEEE Transactions on Knowledge and Data Engineering. 16 (6): 770–773. doi:10.1109/tkde.2004.11. /wiki/Doi_(identifier)
Sigillito, Vincent G., et al. "Classification of radar returns from the ionosphere using neural networks." Johns Hopkins APL Technical Digest10.3 (1989): 262–266.
Zhang, Kun; Fan, Wei (March 2008). "Forecasting skewed biased stochastic ozone days: analyses, solutions and beyond". Knowledge and Information Systems. 14 (3): 299–326. doi:10.1007/s10115-007-0095-1. /wiki/Doi_(identifier)
Reich, Brian J.; Fuentes, Montserrat; Dunson, David B. (March 2011). "Bayesian Spatial Quantile Regression". Journal of the American Statistical Association. 106 (493): 6–20. doi:10.1198/jasa.2010.ap09237. PMC 3583387. PMID 23459794. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3583387
Kohavi, Ron (1996). "Scaling Up the Accuracy of Naive-Bayes Classifiers: A Decision-Tree Hybrid". KDD. 96.
Oza, Nikunj C., and Stuart Russell. "Experimental comparisons of online and batch versions of bagging and boosting." Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, 2001.
Bay, Stephen D. (November 2001). "Multivariate Discretization for Set Mining". Knowledge and Information Systems. 3 (4): 491–512. doi:10.1007/pl00011680. /wiki/Doi_(identifier)
Ruggles, Steven (1995). "Sample designs and sampling errors". Historical Methods. 28 (1): 40–46. doi:10.1080/01615440.1995.9955312. /wiki/Doi_(identifier)
Meek, Christopher, Bo Thiesson, and David Heckerman. "The Learning Curve Method Applied to Clustering." AISTATS. 2001. https://www.microsoft.com/en-us/research/wp-content/uploads/2001/01/lc-aistats.pdf
Fanaee-T, Hadi; Gama, Joao (2013). "Event labeling combining ensemble detectors and background knowledge". Progress in Artificial Intelligence. 2 (2–3): 113–127. doi:10.1007/s13748-013-0040-3. S2CID 3345087. http://repositorio.inesctec.pt/handle/123456789/3506
Giot, Romain; Cherrier, Raphael (2014). "Predicting bikeshare system usage up to one day ahead". 2014 IEEE Symposium on Computational Intelligence in Vehicles and Transportation Systems (CIVTS) (PDF). pp. 22–29. doi:10.1109/CIVTS.2014.7009473. ISBN 978-1-4799-4497-2. 978-1-4799-4497-2
Zhan, Xianyuan; et al. (2013). "Urban link travel time estimation using large-scale taxi data with partial information". Transportation Research Part C: Emerging Technologies. 33: 37–49. Bibcode:2013TRPC...33...37Z. doi:10.1016/j.trc.2013.04.001. /wiki/Bibcode_(identifier)
Moreira-Matias, Luis; et al. (2013). "Predicting taxi–passenger demand using streaming data". IEEE Transactions on Intelligent Transportation Systems. 14 (3): 1393–1402. doi:10.1109/tits.2013.2262376. S2CID 14764358. http://repositorio.inesctec.pt/handle/123456789/5356
Hwang, Ren-Hung; Hsueh, Yu-Ling; Chen, Yu-Ting (2015). "An effective taxi recommender system based on a spatio-temporal factor analysis model". Information Sciences. 314: 28–40. doi:10.1016/j.ins.2015.03.068. /wiki/Doi_(identifier)
H. V. Jagadish, Johannes Gehrke, Alexandros Labrinidis, Yannis Papakonstantinou, Jignesh M. Patel,
Raghu Ramakrishnan, and Cyrus Shahabi. Big data and its technical challenges. Commun. ACM,
57(7):86–94, July 2014.
Caltrans PeMS http://pems.dot.ca.gov/
Meusel, Robert, et al. "The Graph Structure in the Web—Analyzed on Different Aggregation Levels."The Journal of Web Science 1.1 (2015). https://www.nowpublishers.com/article/OpenAccessDownload/JWS-0003
Kushmerick, Nicholas (1999). "Learning to remove Internet advertisements". Proceedings of the third annual conference on Autonomous Agents. pp. 175–181. doi:10.1145/301136.301186. ISBN 978-1-58113-066-9. 978-1-58113-066-9
Fradkin, Dmitriy; Madigan, David (2003). "Experiments with random projections for machine learning". Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining. pp. 517–522. doi:10.1145/956750.956812. ISBN 978-1-58113-737-8. 978-1-58113-737-8
This data was used in the American Statistical Association Statistical Graphics and Computing Sections 1999 Data Exposition.
Ma, Justin; Saul, Lawrence K.; Savage, Stefan; Voelker, Geoffrey M. (2009). "Identifying suspicious URLs: An application of large-scale online learning". Proceedings of the 26th Annual International Conference on Machine Learning. pp. 681–688. doi:10.1145/1553374.1553462. ISBN 978-1-60558-516-1. 978-1-60558-516-1
Levchenko, K.; Pitsillidis, A.; Chachra, N.; Enright, B.; Felegyhazi, M.; Grier, C.; Halvorson, T.; Kanich, C.; Kreibich, C.; He Liu; McCoy, D.; Weaver, N.; Paxson, V.; Voelker, G. M.; Savage, S. (2011). "Click Trajectories: End-to-End Analysis of the Spam Value Chain". 2011 IEEE Symposium on Security and Privacy. pp. 431–446. doi:10.1109/SP.2011.24. ISBN 978-0-7695-4402-1. 978-0-7695-4402-1
Mohammad, Rami M., Fadi Thabtah, and Lee McCluskey. "An assessment of features related to phishing websites using an automated technique."Internet Technology And Secured Transactions, 2012 International Conference for. IEEE, 2012. http://eprints.hud.ac.uk/16229/1/The_7th_ICITST_2012_Conference_-An_Assessment_of_Features_Related_to_Phishing_Websites_using_an_Automated_Technique.pdf
Singh, Ashishkumar; Rumantir, Grace; South, Annie; Bethwaite, Blair (2014). "Clustering Experiments on Big Transaction Data for Market Segmentation". Proceedings of the 2014 International Conference on Big Data Science and Computing. pp. 1–7. doi:10.1145/2640087.2644161. ISBN 978-1-4503-2891-3. 978-1-4503-2891-3
Bollacker, Kurt; Evans, Colin; Paritosh, Praveen; Sturge, Tim; Taylor, Jamie (2008). "Freebase: A collaboratively created graph database for structuring human knowledge". Proceedings of the 2008 ACM SIGMOD international conference on Management of data. pp. 1247–1250. doi:10.1145/1376616.1376746. ISBN 978-1-60558-102-6. 978-1-60558-102-6
Mintz, Mike, et al. "Distant supervision for relation extraction without labeled data." Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2-Volume 2. Association for Computational Linguistics, 2009. https://www.aclweb.org/anthology/P09-1113
Mesterharm, Chris; Pazzani, Michael J. (2011). "Active learning using on-line algorithms". Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining. pp. 850–858. doi:10.1145/2020408.2020553. ISBN 978-1-4503-0813-7. 978-1-4503-0813-7
Wang, Shusen; Zhang, Zhihua (2013). "Improving CUR matrix decomposition and the Nyström approximation via adaptive sampling" (PDF). The Journal of Machine Learning Research. 14 (1): 2729–2769. arXiv:1303.4207. Bibcode:2013arXiv1303.4207W. http://www.jmlr.org/papers/volume14/wang13c/wang13c.pdf
"The Pile". pile.eleuther.ai. Retrieved 14 April 2022. https://pile.eleuther.ai/
"JSON Lines". jsonlines.org. Retrieved 14 April 2022. https://jsonlines.org/
Gao, Leo; Biderman, Stella; Black, Sid; Golding, Laurence; Hoppe, Travis; Foster, Charles; Phang, Jason; He, Horace; Thite, Anish; Nabeshima, Noa; Presser, Shawn (31 December 2020). "The Pile: An 800GB Dataset of Diverse Text for Language Modeling". arXiv:2101.00027 [cs.CL]. /wiki/ArXiv_(identifier)
"The Pile". pile.eleuther.ai. Retrieved 14 April 2022. https://pile.eleuther.ai/
"OSCAR". oscar-project.org. Retrieved 12 August 2023. https://oscar-project.org/
Ortiz Suarez, Pedro, et al. "[2]." Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures. CMLC-7, 2019. https://inria.hal.science/hal-02148693v1/file/Asynchronous_Pipeline_for_Processing_Huge_Corpora_on_Medium_to_Low_Resource_Infrastructures.pdf
Abadji, Julien, et al. "[3]." Towards a Cleaner Document-Oriented Multilingual Crawled Corpus. LREC, 2022. https://aclanthology.org/2022.lrec-1.463.pdf
Cohen, Vanya. "OpenWebTextCorpus". OpenWebTextCorpus. Retrieved 9 January 2023. https://skylion007.github.io/OpenWebTextCorpus/
"openwebtext · Datasets at Hugging Face". huggingface.co. 16 November 2022. Retrieved 9 January 2023. https://huggingface.co/datasets/openwebtext
Saulnier, Lucile (2023). "The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset". arXiv:2303.03915 [cs.CL]. /wiki/ArXiv_(identifier)
"BigScience Data · Datasets at Hugging Face". huggingface.co. 29 August 2023. Retrieved 29 August 2023. https://huggingface.co/bigscience-data
Cattral, Robert; Oppacher, Franz; Deugo, Dwight (2002). "Evolutionary data mining with automatic rule generalization". Recent Advances in Computers, Computing and Communications: 296–300. S2CID 18625415. /wiki/S2CID_(identifier)
Burton, Ariel N.; Kelly, Paul H.J. (August 2006). "Performance prediction of paging workloads using lightweight tracing". Future Generation Computer Systems. 22 (7): 784–793. doi:10.1016/j.future.2006.02.003. /wiki/Doi_(identifier)
Bain, M.; Muggleton, S. (1994). "Learning Optimal Chess Strategies". Machine Intelligence 13. pp. 291–309. doi:10.1093/oso/9780198538509.003.0012. ISBN 978-0-19-853850-9. 978-0-19-853850-9
Quinlan, J. Ross (1983). "Learning Efficient Classification Procedures and Their Application to Chess End Games". Machine Learning. pp. 463–482. doi:10.1007/978-3-662-12405-5_15. ISBN 978-3-662-12407-9. 978-3-662-12407-9
Shapiro, Alen D. (1987). Structured induction in expert systems. Addison-Wesley Longman Publishing Co., Inc.
Matheus, Christopher J.; Rendell, Larry A. (1989). "Constructive Induction on Decision Trees" (PDF). IJCAI. 89. S2CID 11018089. https://www.ijcai.org/Proceedings/89-1/Papers/103.pdf
Belsley, David A., Edwin Kuh, and Roy E. Welsch. Regression diagnostics: Identifying influential data and sources of collinearity. Vol. 571. John Wiley & Sons, 2005.
Ruotsalo, Tuukka; Aroyo, Lora; Schreiber, Guus (2009). "Knowledge-based linguistic annotation of digital cultural heritage collections" (PDF). IEEE Intelligent Systems. 24 (2): 64–75. doi:10.1109/MIS.2009.32. hdl:1871.1/9f6091aa-9596-46a9-9251-f11edeeb28b7. S2CID 6667472. Archived from the original (PDF) on 16 August 2017. Retrieved 6 December 2018. https://web.archive.org/web/20170816023938/http://dare.ubvu.vu.nl/bitstream/handle/1871/24407/243319.pdf?sequence=3
Li, Lihong; Chu, Wei; Langford, John; Wang, Xuanhui (2011). "Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms". Proceedings of the fourth ACM international conference on Web search and data mining. pp. 297–306. arXiv:1003.5956. doi:10.1145/1935826.1935878. ISBN 978-1-4503-0493-1. 978-1-4503-0493-1
Yeung, Kam Fung; Yang, Yanyan (2010). "A Proactive Personalized Mobile News Recommendation System". 2010 Developments in E-systems Engineering. pp. 207–212. doi:10.1109/DeSE.2010.40. ISBN 978-1-4244-8044-9. 978-1-4244-8044-9
Gass, Susan E.; Roberts, J. Murray (2006). "The occurrence of the cold-water coral Lophelia pertusa (Scleractinia) on oil and gas platforms in the North Sea: colony growth, recruitment and environmental controls on distribution". Marine Pollution Bulletin. 52 (5): 549–559. Bibcode:2006MarPB..52..549G. doi:10.1016/j.marpolbul.2005.10.002. PMID 16300800. /wiki/Bibcode_(identifier)
Gionis, Aristides; Mannila, Heikki; Tsaparas, Panayiotis (March 2007). "Clustering aggregation". ACM Transactions on Knowledge Discovery from Data. 1 (1): 4. doi:10.1145/1217299.1217303. /wiki/Doi_(identifier)
Obradovic, Zoran, and Slobodan Vucetic.Challenges in Scientific Data Mining: Heterogeneous, Biased, and Large Samples. Technical Report, Center for Information Science and Technology Temple University, 2004.
Van Der Putten, Peter; van Someren, Maarten (2000). "CoIL challenge 2000: The insurance company case". Published by Sentient Machine Research, Amsterdam. Also a Leiden Institute of Advanced Computer Science Technical Report. 9: 1–43.
Mao, K. Z. (2002). "RBF neural network center selection based on Fisher ratio class separability measure". IEEE Transactions on Neural Networks. 13 (5): 1211–1217. doi:10.1109/tnn.2002.1031953. PMID 18244518. /wiki/Doi_(identifier)
Olave, Manuel; Rajkovic, Vladislav; Bohanec, Marko (1989). "An application for admission in public school systems" (PDF). Expert Systems in Public Administration. 1: 145–160. http://kt.ijs.si/MarkoBohanec/pub/Nursery89.pdf
Lizotte, Daniel J.; Madani, Omid; Greiner, Russell (2012). "Budgeted Learning of Naive-Bayes Classifiers". arXiv:1212.2472 [cs.LG]. /wiki/ArXiv_(identifier)
Lebowitz, Michael (1984). Concept Learning in a Rich Input Domain: Generalization-Based Memory (Report). doi:10.7916/D8KP8990. /wiki/Doi_(identifier)
Yeh, I-Cheng; Yang, King-Jang; Ting, Tao-Ming (2009). "Knowledge discovery on RFM model using Bernoulli sequence". Expert Systems with Applications. 36 (3): 5866–5871. doi:10.1016/j.eswa.2008.07.018. /wiki/Doi_(identifier)
Lee, Wen-Chen; Cheng, Bor-Wen (2011). "An intelligent system for improving performance of blood donation". Journal of Quality Vol. 18 (2): 173. http://www.airitilibrary.com/Publication/alDetailedMesh?docid=10220690-201104-201105050019-201105050019-173-185
Schmidtmann, Irene, et al. "Evaluation des Krebsregisters NRW Schwerpunkt Record Linkage Archived 6 December 2018 at the Wayback Machine." Abschlußbericht vom 11 (2009). http://www.krebsregister-nrw.de/fileadmin/user_upload/dokumente/Evaluation/EKR_NRW_Evaluation_Abschlussbericht_2009-06-11.pdf
Sariyar, Murat; Borg, Andreas; Pommerening, Klaus (2011). "Controlling false match rates in record linkage using extreme value theory". Journal of Biomedical Informatics. 44 (4): 648–654. doi:10.1016/j.jbi.2011.02.008. PMID 21352952. /wiki/Doi_(identifier)
Candillier, Laurent; Lemaire, Vincent (August 2013). "Active learning in the real-world design and analysis of the Nomao challenge". The 2013 International Joint Conference on Neural Networks (IJCNN). Vol. 8. pp. 1–8. doi:10.1109/IJCNN.2013.6706908. ISBN 978-1-4673-6129-3. 978-1-4673-6129-3
Garrido Marquez, Ivan (2013). A domain adaptation method for text classification based on self-adjusted training approach (Thesis).[page needed] https://inaoe.repositorioinstitucional.mx/jspui/handle/1009/230
Nagesh, Harsha S., Sanjay Goil, and Alok N. Choudhary. "Adaptive Grids for Clustering Massive Data Sets." SDM. 2001.
Kuzilek, Jakub, et al. "OU Analyse: analysing at-risk students at The Open University." Learning Analytics Review (2015): 1–16. http://oro.open.ac.uk/42529/1/__userdata_documents4_ctb44_Desktop_analysing-at-risk-students-at-open-university.pdf
Siemens, George, et al. Open Learning Analytics: an integrated & modularized platform. Diss. Open University Press, 2011. https://www.solaresearch.org/core/open-learning-analytics-an-integrated-modularized-platform/
Barlacchi, Gianni; De Nadai, Marco; Larcher, Roberto; Casella, Antonio; Chitic, Cristiana; Torrisi, Giovanni; Antonelli, Fabrizio; Vespignani, Alessandro; Pentland, Alex; Lepri, Bruno (27 October 2015). "A multi-source dataset of urban life in the city of Milan and the Province of Trentino". Scientific Data. 2 (1). Bibcode:2015NatSD...250055B. doi:10.1038/sdata.2015.55. PMC 4622222. PMID 26528394. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4622222
Vanschoren J, van Rijn JN, Bischl B, Torgo L (2013). "OpenML: networked science in machine learning". SIGKDD Explorations. 15 (2): 49–60. arXiv:1407.7722. doi:10.1145/2641190.2641198. S2CID 4977460. /wiki/ArXiv_(identifier)
Olson RS, La Cava W, Orzechowski P, Urbanowicz RJ, Moore JH (2017). "PMLB: a large benchmark suite for machine learning evaluation and comparison". BioData Mining. 10 (1): 36. arXiv:1703.00512. Bibcode:2017arXiv170300512O. doi:10.1186/s13040-017-0154-4. PMC 5725843. PMID 29238404. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5725843
"Off The Shelf Datasets". appen.com. Appen. Retrieved 30 December 2020. https://appen.com/off-the-shelf-datasets/
"Open Source Datasets". appen.com. Appen. Retrieved 30 December 2020. https://appen.com/resources/datasets/