We describe L-SR1, a new second order method to train deep neural networks. Second order methods hold great promise for distributed training of deep networks. Unfortunately, they have not proven practical. Two significant barriers to their success are inappropriate handling of saddle points, and poor conditioning of the Hessian. L-SR1 is a practical second order method that addresses these concerns. We provide preliminary experimental results showing that L-SR1 performs at least as well as several other first order methods, and better than L-BFGS, on the MNIST and CIFAR10 datasets.