Every paper below is cited by at least one note in this garden, and the back-links say which. Titles, author lists, and years were read off each publisher’s own page rather than from memory, so the citations here match the record. Where a note leans on a paper for a specific claim, that note’s own ## Sources section says which claim.

Architectures and representation learning

Attention Is All You Need (2017), Ashish Vaswani, Noam Shazeer, Niki Parmar et al. The transformer. Cited by Attention and Transformers.

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018), Jacob Devlin, Ming-Wei Chang, Kenton Lee et al. Cited by Attention and Transformers.

Deep Residual Learning for Image Recognition (2015), Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. Residual connections. Cited by Pooling and CNN Architectures.

U-Net: Convolutional Networks for Biomedical Image Segmentation (2015), Olaf Ronneberger, Philipp Fischer, Thomas Brox. Cited by Diffusion Models.

Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation (2014), Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre et al. The GRU. Cited by Recurrent Neural Networks.

Efficient Estimation of Word Representations in Vector Space (2013), Tomas Mikolov, Kai Chen, Greg Corrado et al. word2vec. Cited by Embeddings.

Distributed Representations of Words and Phrases and their Compositionality (2013), Tomas Mikolov, Ilya Sutskever, Kai Chen et al. Cited by Embeddings.

node2vec: Scalable Feature Learning for Networks (2016), Aditya Grover, Jure Leskovec. Cited by Embeddings.

Generative models

Generative Adversarial Networks (2014), Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza et al. Cited by GANs.

Auto-Encoding Variational Bayes (2013), Diederik P Kingma, Max Welling. The VAE. Cited by Autoencoders.

Denoising Diffusion Probabilistic Models (2020), Jonathan Ho, Ajay Jain, Pieter Abbeel. Cited by Diffusion Models.

Training, optimization, and initialization

Adam: A Method for Stochastic Optimization (2014), Diederik P. Kingma, Jimmy Ba. Cited by Faster Optimizers and LR Scheduling, Gradient Descent.

An overview of gradient descent optimization algorithms (2016), Sebastian Ruder. Cited by Gradient Descent.

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift (2015), Sergey Ioffe, Christian Szegedy. Cited by Regularization in Deep Learning, Vanishing and Exploding Gradients.

Understanding the difficulty of training deep feedforward neural networks (2010), Xavier Glorot, Yoshua Bengio. Xavier initialization. Cited by Activation Functions, Vanishing and Exploding Gradients.

Deep Sparse Rectifier Neural Networks (2011), Xavier Glorot, Antoine Bordes, Yoshua Bengio. The ReLU case. Cited by Activation Functions.

Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification (2015), Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. He initialization and PReLU. Cited by Vanishing and Exploding Gradients.

On the difficulty of training Recurrent Neural Networks (2012), Razvan Pascanu, Tomas Mikolov, Yoshua Bengio. Gradient clipping. Cited by Vanishing and Exploding Gradients.

On the importance of initialization and momentum in deep learning (2013), Ilya Sutskever, James Martens, George Dahl et al. Cited by Faster Optimizers and LR Scheduling.

Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates (2017), Leslie N. Smith, Nicholay Topin. Cited by Faster Optimizers and LR Scheduling.

XGBoost: A Scalable Tree Boosting System (2016), Tianqi Chen, Carlos Guestrin. Cited by Decision Trees and Ensembles.

Transfer, meta-learning, and interpretability

How transferable are features in deep neural networks? (2014), Jason Yosinski, Jeff Clune, Yoshua Bengio et al. Cited by Transfer Learning.

Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (2017), Chelsea Finn, Pieter Abbeel, Sergey Levine. MAML. Cited by Meta-Learning.

Meta-Learning in Neural Networks: A Survey (2020), Timothy Hospedales, Antreas Antoniou, Paul Micaelli et al. Cited by Meta-Learning.

Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps (2013), Karen Simonyan, Andrea Vedaldi, Andrew Zisserman. Cited by Feature Attribution and Saliency.

Axiomatic Attribution for Deep Networks (2017), Mukund Sundararajan, Ankur Taly, Qiqi Yan. Integrated gradients. Cited by Feature Attribution and Saliency.

Reinforcement learning

Playing Atari with Deep Reinforcement Learning (2013), Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al. Cited by Deep Reinforcement Learning.

Human-level control through deep reinforcement learning (2015), Volodymyr Mnih, Koray Kavukcuoglu, David Silver et al., Nature. DQN. Cited by Deep Reinforcement Learning.

Mastering the game of Go with deep neural networks and tree search (2016), David Silver, Aja Huang, Chris J. Maddison et al., Nature. AlphaGo. Cited by Deep Reinforcement Learning.

Applications and fairness

Unsupervised word embeddings capture latent knowledge from materials science literature (2019), Vahe Tshitoyan, John Dagdelen, Leigh Weston et al., Nature. Cited by Embeddings.

Inherent Trade-Offs in the Fair Determination of Risk Scores (2016), Jon Kleinberg, Sendhil Mullainathan, Manish Raghavan. Cited by The Impossibility of Algorithmic Fairness.

Fair prediction with disparate impact: A study of bias in recidivism prediction instruments (2017), Alexandra Chouldechova. Cited by The Impossibility of Algorithmic Fairness.

People are not coins. Morally distinct types of predictions necessitate different fairness constraints (2022), Eleonora Vigano’, Corinna Hertweck, Christoph Heitz et al. Cited by The Impossibility of Algorithmic Fairness.

What’s Sex Got To Do With Fair Machine Learning? (2020), Lily Hu, Issa Kohler-Hausmann. Cited by Social Categories and Machine Learning.

Could a Large Language Model be Conscious? (2023), David J. Chalmers. Cited by Could an LLM Be Conscious?.

Security

The Protection of Information in Computer Systems, Jerome H. Saltzer and Michael D. Schroeder. The invited paper that gave the field least privilege and the rest of its design principles. Cited by Privilege Separation and Least Privilege.

A Future-Adaptable Password Scheme (1999), Niels Provos and David Mazieres, FREENIX Track, USENIX Annual Technical Conference. bcrypt. Cited by Password Hashing and Salting.

Preventing Privilege Escalation (2003), Niels Provos, Markus Friedl, and Peter Honeyman, 12th USENIX Security Symposium. Privilege separation in OpenSSH. Cited by Privilege Separation and Least Privilege.

Philosophy of mind and computation

Peer-reviewed entries from the Stanford Encyclopedia of Philosophy, which the ethics and language-theory notes lean on.

The Lambda Calculus, Jesse Alama and Johannes Korbmacher. Cited by Lambda Calculus: Syntax and Substitution.

Consciousness (2004), Robert Van Gulick. Cited by Consciousness: Access vs Phenomenal.

Higher-Order Theories of Consciousness (2001), Peter Carruthers and Rocco Gennaro. Cited by Scientific Theories of Consciousness.

Functionalism (2004), Janet Levin. Cited by Functionalism and Multiple Realizability, The Biological Substrate Objection.

Multiple Realizability (1998), John Bickle. Cited by Functionalism and Multiple Realizability.

Physicalism (2001), Daniel Stoljar. Cited by Physicalism and the Mind.

Moral Responsibility (2019), Matthew Talbert. Cited by Fairness as Equal Concern.

Sources