Published onApril 16, 2026SPD 1: Analysis of component activations of deep toy model of superpositionAnalysing patterns of component activations on 2-6L tied and untied SPD-based decomposed TMS models.
Published onFebruary 22, 2026SNMF 1: How does the type of activations and number of top activating tokens affect concept detection in LLMs?A small reproduction and extension of the concept detection results of (Shafran et al., 2025) paper.
Published onJanuary 16, 2026Reproduction of the KEEN paper on estimation of knowledge on QA datasetA small reproduction of the main thesis of the KEEN (Gottesman & Geva, 2024) paper.
Published onNovember 12, 2025Listwise loss does not consistently makeup for retrieval performance drop induced by InfoNCE lossReplicating Tamber et al. with Nomic-Embed and ModernBERT-Embed shows that conventional contrastive learning can fail, but not uniformly across embedding models.