| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuseNote that this will soon get outdated. To know more about me and my research, a much better place is my website: 7vik.io.
Independent Researcher | AGI Safety • Interpretability • Reinforcement Learning
📍 UC Berkeley (CHAI), MATS Program
📧 zsatvik@gmail.com | 🌐 7vik.io | Scholar | LinkedIn | GitHub
On a quest to understand intelligence and ensure that advanced AGI is safe and beneficial.
I’m an independent AI safety researcher currently working with:
Previously:
=Equal contribution; full list at Google Scholar
Intricacies of Feature Geometry in Large Language Models
ICLR 2025 (poster); Runner-up, ICLR Blog Awards
Code | Blog
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
Under Review
Poster | Blog
Auditing Language Models for Hidden Objectives
Anthropic (external collaboration)
Anthropic Blog | Blog
Progress Measures for Grokking on Real-world Tasks
ICML 2024 Workshop on High-dimensional Learning Dynamics
Code
Challenges in Mechanistically Interpreting Model Representations
ICML 2024 Workshop on Mechanistic Interpretability
Code
A is for Absorption: Studying Feature Splitting and Absorption in SAEs
Under Review
CataractBot: An LLM-Powered Expert-in-the-Loop Chat System
IMWUT / UbiComp 2025
Code
Predicting Treatment Adherence of Tuberculosis Patients at Scale
PMLR 2022; Outstanding Paper, NeurIPS 2022
Media Coverage
NICE: Normalized Invariance to Choice of Example
This repository holds a list of cool resources for Silica.
| Back | FazBrowse Home | New Git URL |