Showing posts with label neural networks. Show all posts
Showing posts with label neural networks. Show all posts

Tuesday, December 1, 2020

The Information Theory of Developmental Pruning: Optimizing Global Network Architecture Using Local Synaptic Rules

Another paper from the Hennig lab is out, this one is from Carolin Scholl's master's thesis. Once again, we used an artificial neural network to get intuition about biology. The paper is on BioRiv, and you can also get the PDF here. 

Sunday, February 10, 2019

Neural field models for latent state inference

The final paper from my Edinburgh postdoc in the Sanguinetti and Hennig labs (perhaps, we shall see). 

[get PDF]

We combined neural field modelling with point-process latent state inference. Neural field models capture collective population activity like oscillations and spatiotemporal waves. They make the simplifying assumption that neural activity can be summarized by the average firing rate in a region.

High-density electrode array recordings can now record developmental retinal waves in detail. We derived a neural field model for these waves from the microscopic model proposed by Hennig et al.. This model posits that retinal waves are supported by an quiescent, active, and refractory states. 

Fig 3. Spatial 3-state neural-field model exhibits self-organized multi-scale wave phenomena. Simulated example states at selected time-points on a [0,1]² unit interval using a 20×20 grid with effective population density of $\rho{=}50$ cells per unit area, and rate parameters $\sigma{=}0.075$, $\rho_a {=} 0.4$, $\rho_r {=} 3.2 \times 10^{−3}$, $\rho_e {=} 0.028$, and $\rho_q {=} 0.25$ (Methods: Sampling from the model). As, for instance, in neonatal retinal waves, spontaneous excitation of quiescent cells (blue) lead to propagating waves of activity (red), which establish localized patches in which cells are refractory (green) to subsequent wave propagation. Over time, this leads to diverse patterns of waves at a range of spatial scales. 

Friday, July 13, 2018

Local learning rules to attenuate forgetting in neural networks

Another paper from our work on Restricted Boltzmann Machines (RBMs) from the Hennig lab. 

[get paper PDF]

Main points:
  • We noticed that measures of synaptic importance were available from local firing statistics (at least in Boltzmann machines)
  • We look at an artificial neural network that stores memories and is easy to analyze. (Hopfield nets are the zero temperature limit of a Boltzmann machine).
  • We evaluated whether this local measure of synaptic importance could help stabilize important weights when networks learn multiple things that interfere with each-other
  • Intuition: biological variables, like synapse size, can correlate with useful statistical quantities. This provides tricks for biologically-plausible approximations of algorithms.
  • Intuition: in systems that learn, if a parameter takes on an unusual or surprising value, it is likely that this value was set through learning—and you might want to leave it fixed.

Wednesday, February 28, 2018

Optimal encoding in stochastic latent-variable models

Update: the review process for this was a long one, But! It is published now, in Entropy [PDF].

The sensory system encodes the external world using spiking population codes, and must contend with the fixed bandwidth and noise inherent to neural communication. In this work, we explored a machine-learning model that shares similar constraints, and examined the coding strategies it learned when trained to encode visual input.

We used Restricted Boltzmann Machines (RBMs), which are a two-layer stochastic binary neural network that learn generative models of their inputs. We explored how optimized encoders handle limited encoding resources by varying the number of binary "neurons" available to represent visual inputs.

We found several statistical signatures that emerge around an "optimal" model size: one that is just large enough to encode its inputs. Around the optimal model model size, we saw emergence of statistical features often observed in neural population codes: sparsity, decorrelation, statistical criticality, and variability suppression.

We interpret the learned encoding strategies as a way to encode stimuli with variable bit-rates over a noisy channel of fixed bandwidth. Low-noise regions of coding space are reserved for stimuli that take the most bits to describe. In our simulations, the networks learned to suppress noise by strongly silencing some neurons.

Common stimuli don't require very many bits to transmit (from a Shannon coding perspective). Such stimuli are encoded in the noisier regions of the neural code, and exhibit increased variability. Increase neural variability in the presence of limited information is another feature observed in neural population codes.

Preview: Figure 5

 

Analyses of parameter sensitivity suggests an optimal model size for encoding sensory statistics: (a) Analysis of the Fisher Information Matrix (FIM) over a range of hidden-layer sizes (top to bottom; 13 visible units). From left to right, (1) FIM eigenvalue spectra $\lambda_i$ (y-axis) over a range of inverse temperatures β indicate that model fits (β=1) past a certain size lie at a peak in their generalized susceptibility. This is a correlate of criticality in Ising spin models. Eigenvalues below $10^{−5}$ are truncated, and the largest and smallest eigenvalues are in red; (2) Important parameters in the leading FIM eigenvector align with individual hidden units, and become sparse for larger hidden layers. The eigenvector is displayed separately for the weights (matrix), and the visible (vertical) and hidden (horizontal) biases; (3) The average sensitivity of each parameter over all FIM eigenvectors, shown here as the square root of the FIM diagonal, also shows sparsity, indicating that beyond a certain size additional hidden units contribute little to model accuracy. Data is shown as in column 2; (4) Variance of the hidden unit activation as a function of stimulus energy. In larger models, units with sensitive parameters contribute to encoding low energy, less informative patterns. (b) The average sensitivity of each parameter, measured by the trace of the FIM, normalized by hidden-layer size, decreases as hidden-layer size grows. (c) Hidden unit projective fields from a model with 37 visible and 60 hidden units, ordered by relative sensitivity (rank indicated above each image). More important units (ranks 1–8) encode spatially simple features such as localized patches, while the least important ones (ranks 53–60) have complex features.

Overall, 

We showed that machine learning models that resemble spiking population codes learn coding strategies tantalizingly similar to what is seen in vivo. We also learned that the statistical features of the population code can be used to check whether a model is too small, too large, or just the right size, to encode its inputs. 

Many thanks to Martino Sorbaro and Matthias H. Hennig for sticking with this paper during the long review. This work can be cited as

Rule, M.E., Sorbaro, M. and Hennig, M.H., 2020. Optimal encoding in stochastic latent-variable Models. Entropy, 22(7), p.714.

Monday, January 1, 2018

Autoregressive point-processes as latent state-space models

In 2016 I started a postdoc with the labs of Guido Sanguinetti and Matthias H. Hennig—and this is the first paper to result!

[get PDF]

A central challenge in neuroscience is understanding how the activity of single cells combines to create the collective dynamics that underlie perception, cognition, and behavior. One way to study this is to build detailed models, and "coarse grain" them to see which details are important. 

Our paper develops ways to relate detailed point-process models to coarse-grained quantities, like average neuronal firing rates and correlations. Point-process models are used for statistical modelling of spike train data. They can reveal effective neural dynamics by capturing how neurons in a population inhibit or excite each-other and themselves. 

Preview of figures: 

Figure 2: Moment closure of autoregressive PPGLMs combines aspects of three modeling approaches:

(A) Log-linear autoregressive PPGLM framework (e.g., Weber & Pillow, 2017). Dependence on the history of both extrinsic covariates x(t) and the process itself y(t) are mediated by linear filters, which are combined to predict the instantaneous log intensity of the process. (B) Latent state-space models learn a hidden dynamical system, which can be driven by both extrinsic covariates and spiking outputs. Such models are often fit using expectation- maximization, and the learned dynamics are descriptive. (C) Moment closure recasts autoregressive PPGLMs as state-space models. History dependence of the process is subsumed into the state-space dynamics, but the latent states retain a physical interpretation as moments of the process history (dashed arrow). (D) Compare to neural mass and neural field models, which define dynamics on a state space with a physical interpretation as moments of neural population activity

Friday, September 1, 2017

Population coding of sensory stimuli through latent variables

Edit: This work is published now, in Entropy [PDF].

Martino Sorbaro has been doing some really interesting work exploring the encoding strategies learned by artificial neural networks. We've found similarities between the statistics of the population codes learned by Restricts Boltzmann Machines (RBMs), and those of the retina. We'll present this work as a poster at the upcoming  Integrated Systems Neuroscience in Manchester.

TL;DR:

RBMs as a model for latent-variable encoding

  • Optimal latent-variable encoding of visual stimuli seems to consistently yield models near statistical criticality.  
  • Poor fits (too few hidden units,under-fitting) do not exhibit this property.
  • Critical RBMs mimic the retina in Zipf laws, sparsity, and decorrelation.
  • Above the optimal model size, extra units are weakly constrained as measured by Fisher information.  
  • Receptive fields of excess units are less retina-like.

Questions and controversy

  • Is statistical criticality a general feature of factorized latent variable models?
  • Is criticality in the retina expected based simply on optimal encoding?

[download poster PDF]


Abtract:

Several studies observe power-law statistics consistent with critical scaling exponents in neural data, but it is unclear whether such statistics necessarily imply criticality. In this work, we examine whether the 1/f statistics of retinal populations are inherited from visual stimuli, or whether they might emerge from collective neural dynamics independently of stimulus statistics. We examine, in silico, a latent-variable encoding model of visual scenes, and empirically explore the conditions under which such a model exhibits 1/f statistics thought to reflect criticality. Specifically, we examine the Restricted Boltzmann Machines (RBMs) as a factorized binary latent-variable model for stimulus encoding. We find two surprising results. First, latent variable models need not exhibit 1/f statistics, but that the optimal model size, reflecting the smallest model that can faithfully encode stimuli, does. We illustrate that the optimal model size can be predicted from sloppy dimensions of the Fisher information matrix (FIM), which align with a subspace spanning the superfluous latent variables. Second, the optimal-sized model can exhibit 1/f statistics even when stimuli do not, indicating that this property is not inherited from environmental statistics. Furthermore, such models exhibit properties of statistical criticality, including diverging susceptibilities. This empirical evidence suggests that 1/f statistics are neither inherited from the  environment, nor a necessary feature of accurate encoding. Rather, it suggests that parsimonious latent- variable models are naturally poised close to criticality, generating the observed 1/f statistics. Overall,  these results are consistent with conjectures in other fields that a cost-benefit trade-off between expressivity and parsimony underlies the emergence of criticality and 1/f power-law statistics. Furthermore, this works suggests that in latent-variable encoding models, the emergence of 1/f statistics reflects true criticality and is not inherited from the environmental distribution of stimuli.

The poster can be cited as:

Sorbaro, M, Rule, M., Hilgen, G., Sernagot, E. , D, Hennig, M. H. (2017) Signatures of optimal population coding of sensory stimuli through latent variables. [Poster] The second Integrated Systems Neuroscience Workshop, 7-8th September 2017, at The University of Manchester, Manchester, UK.

Edit: the paper can be cited as

Rule, M.E., Sorbaro, M. and Hennig, M.H., 2020. Optimal encoding in stochastic latent-variable Models. Entropy, 22(7), p.714.

 

Sunday, March 8, 2009

Self-organizing maps

I'm currently following a course at CMU on neural networks. This post explores learning a 2D embedding of a complex perceptual spacing using self-organizing maps. These outputs were computed using the Lightweight Efficient Network Simulator. The learned embedding makes it possible to wander randomly through the latent low-dimensional manifold underlying the structure in high-dimensional data, e.g. human poses