kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Online reinforcement learning of state representation in recurrent network supported by the power of random feedback and biological constraints
Univ Tokyo, Grad Sch Educ, Phys & Hlth Educ, Tokyo, Japan; Okinawa Inst Sci & Technol, Theoret Sci Visiting Program, Okinawa, Japan.
Okinawa Inst Sci & Technol, Theoret Sci Visiting Program, Okinawa, Japan; Icahn Sch Med Mt Sinai, Dept Psychiat, New York, NY USA.
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Computational Science and Technology (CST). KTH, Centres, Science for Life Laboratory, SciLifeLab. Okinawa Inst Sci & Technol, Theoret Sci Visiting Program, Okinawa, Japan.ORCID iD: 0000-0002-8044-9195
Univ Tokyo, Grad Sch Educ, Phys & Hlth Educ, Tokyo, Japan; Okinawa Inst Sci & Technol, Theoret Sci Visiting Program, Okinawa, Japan; Univ Tokyo, Int Res Ctr Neurointelligence WPI IRCN, Tokyo, Japan.
2025 (English)In: eLIFE, E-ISSN 2050-084X, Vol. 14, article id RP104101Article in journal (Refereed) Published
Abstract [en]

Representation of external and internal states in the brain plays a critical role in enabling suitable behavior. Recent studies suggest that state representation and state value can be simultaneously learned through Temporal-Difference-Reinforcement-Learning (TDRL) and Backpropagation-Through-Time (BPTT) in recurrent neural networks (RNNs) and their readout. However, neural implementation of such learning remains unclear as BPTT requires offline update using transported downstream weights, which is suggested to be biologically implausible. We demonstrate that simple online training of RNNs using TD reward prediction error and random feedback, without additional memory or eligibility trace, can still learn the structure of tasks with cue–reward delay and timing variability. This is because TD learning itself is a solution for temporal credit assignment, and feedback alignment, a mechanism originally proposed for supervised learning, enables gradient approximation without weight transport. Furthermore, we show that biologically constraining downstream weights and random feedback to be non-negative not only preserves learning but may even enhance it because the non-negative constraint ensures loose alignment—allowing the downstream and feedback weights to roughly align from the beginning. These results provide insights into the neural mechanisms underlying the learning of state representation and value, highlighting the potential of random feedback and biological constraints.

Place, publisher, year, edition, pages
eLife Sciences Publications, Ltd , 2025. Vol. 14, article id RP104101
Keywords [en]
dopamine, corticostriatal, reinforcement learning, state representation, feedback alignment, biological constraints
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:kth:diva-374577DOI: 10.7554/eLife.104101ISI: 001579093100001PubMedID: 40991326OAI: oai:DiVA.org:kth-374577DiVA, id: diva2:2023228
Note

QC 20251218

Available from: 2025-12-18 Created: 2025-12-18 Last updated: 2025-12-18Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMed

Authority records

Kumar, Arvind

Search in DiVA

By author/editor
Kumar, Arvind
By organisation
Computational Science and Technology (CST)Science for Life Laboratory, SciLifeLab
In the same journal
eLIFE
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
urn-nbn

Altmetric score

doi
pubmed
urn-nbn
Total: 28 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf