Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

To expand on this a bit (please correct me if I'm wrong), there are three major uses of dimensionality reduction:

1. Reduce overfitting

2. Train models faster

3. Take advantage of unsupervised data

Regularization handles case (1) quite well. It can also be used in conjunction with most methods of dimensionality reduction such as PCA/auto-encoders.

PCA covers all three, but isn't as effective at dimensionality reduction compared to auto-encoders.

Auto-encoders tend to yield a better compression than PCA but take more time to train and produce output that's harder to understand. There is a bit of analogy here, auto-encoders are to PCA what neural nets are to linear regression.



L1 regularization can also be used for feature selection[0] and hence dimensionality reduction. I've in found practice that the speed of this can compare favorably even to fast implementations[1] of svd.

[0] Feature selection, L1 vs. L2 regularization, and rotational invariance, Andrew Y. Ng. In Proceedings of the Twenty-first International Conference on Machine Learning, 2004. http://ai.stanford.edu/~ang/papers/icml04-l1l2.pdf

[1]"Augmented Implicitly Restarted Lanczos Bidiagonalization Methods", J. Baglama and L. Reichel, SIAM J. Sci. Comput. 2005.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: