Behavior of Chain-of-thought Monitorability
Studying how monitor effectiveness changes as the capability gap between monitor and target models widens, with a case study on distinguishing sandbagging from genuine incapability.
A collection of projects I've worked on, ranging from academic research to personal explorations. If you think it would be useful or interesting to collaborate on a project, please contact me to discuss.
Studying how monitor effectiveness changes as the capability gap between monitor and target models widens, with a case study on distinguishing sandbagging from genuine incapability.
Extended Tau2-Bench for low-resource Southeast Asian languages and localized agentic evaluation.
AI chatbot generating memorable English and Mandarin mnemonics with QLoRA fine-tuning and DPO preference modeling.
Minimal Llama-2 implementation in PyTorch with RoPE, RMSNorm, SwiGLU, self-attention, and transformer blocks.
Contributed to RL runtime environments and community-based AI benchmarks.
Analysis of LLM factuality, hallucination, topic patterns, and cross-cultural reasoning failures on multicultural riddles.
Designed the SEACrowd website and managed social content for a Southeast Asian AI research community.
Created and maintain a personal academic Astro theme for research portfolios, publications, projects, and blogs.
Evaluated whether implicit values in agentic LLMs are inclusive of animals, by extending CaML's The Animal Compassion
Investigated multilingual LLM representations with SAEs and feature steering across 67 languages.
AI-powered Hex dashboard scoring safety across 41 San Francisco neighborhoods with live DataSF and ClickHouse data.
Luma-like platform for discovering, managing, and RSVPing to local recreational sport events.
Full-stack PWA for seizure tracking and predictive warnings, with React, Flask, XGBoost, and LSTM models.
Question-answering system over Obsidian-style personal notes using LlamaIndex, OpenAI API, and retrieval-augmented generation.
Replicated and extended a synthetic control analysis of Philadelphia SNAP benefit redemption in R.
Analyze GP visit patterns using Zero-Inflated Poisson models with complete and partial pooling, with Bayesian inference and data imputation.
Text classification system to automatically classify notes and assignments using SVM and Naive Bayes with 87% accuracy.
My GitHub ↗ contributions over the past year. Colored squares represent days with commits.