Projects

A collection of projects I've worked on, ranging from academic research to personal explorations. If you think it would be useful or interesting to collaborate on a project, please contact me to discuss.

Behavior of Chain-of-thought Monitorability

- Present

Studying how monitor effectiveness changes as the capability gap between monitor and target models widens, with a case study on distinguishing sandbagging from genuine incapability.

Chain of Thought Sandbagging Inspect AI OpenRouter API

SEATauBench

-

Extended Tau2-Bench for low-resource Southeast Asian languages and localized agentic evaluation.

LLM Evaluation Agentic AI Southeast Asian Languages Localization SEACrowd

Mnemonic Generation for Vocabulary Learning

-

AI chatbot generating memorable English and Mandarin mnemonics with QLoRA fine-tuning and DPO preference modeling.

unsloth QLoRA DPO trl Gemma DeepSeek

Mini-LLaMA2 PyTorch Implementation

Minimal Llama-2 implementation in PyTorch with RoPE, RMSNorm, SwiGLU, self-attention, and transformer blocks.

Large Language Model (LLM) PyTorch Transformers

Other Projects

BenchFlow (Open Source)

- Present

Contributed to RL runtime environments and community-based AI benchmarks.

Open Source Benchmarking harbor AI Agents

Multicultural Riddles Benchmarking

- Present

Analysis of LLM factuality, hallucination, topic patterns, and cross-cultural reasoning failures on multicultural riddles.

LLM Evaluation Multicultural AI Hallucination Cohere Labs

SEACrowd Website (Design & Growth)

- Present

Designed the SEACrowd website and managed social content for a Southeast Asian AI research community.

Jekyll Bootstrap JavaScript SCSS

Astro Scholar Theme for Academics

- Present

Created and maintain a personal academic Astro theme for research portfolios, publications, projects, and blogs.

Astro TypeScript Academic Portfolio Satteri

Animals-Aligned Agentic AI Benchmarking

-

Evaluated whether implicit values in agentic LLMs are inclusive of animals, by extending CaML's The Animal Compassion

AI Safety Agentic AI Animal Welfare Benchmarking Futurekind

Multilingual Representations with Sparse Autoencoders

Investigated multilingual LLM representations with SAEs and feature steering across 67 languages.

Sparse Autoencoders Mechanistic Interpretability Multilingual AI HuggingFace LaBSE

San Francisco Safety Dashboard

AI-powered Hex dashboard scoring safety across 41 San Francisco neighborhoods with live DataSF and ClickHouse data.

Hex ClickHouse DataSF Dashboard

SportConnect

Luma-like platform for discovering, managing, and RSVPing to local recreational sport events.

TypeScript React Express PostgreSQL TailwindCSS DaisyUI BetterAuth

SeizureSavvy

-

Full-stack PWA for seizure tracking and predictive warnings, with React, Flask, XGBoost, and LSTM models.

Progressive Web App (PWA) React Flask XGBoost LSTM SQLAlchemy

Ask My Second Brain with RAG

-

Question-answering system over Obsidian-style personal notes using LlamaIndex, OpenAI API, and retrieval-augmented generation.

Retrieval Augmented Generation (RAG) Python LlamaIndex OpenAI

Causal Inference of Political Intervention

Replicated and extended a synthetic control analysis of Philadelphia SNAP benefit redemption in R.

R rmarkdown Causal Inference Synthetic Control Replication

Bayesian Hierarchical Modeling for GP Visit Count Data

Analyze GP visit patterns using Zero-Inflated Poisson models with complete and partial pooling, with Bayesian inference and data imputation.

bayesian modeling hierarchical models PyMC missing data imputation healthcare analytics statistical inference

Text classification using SVM and Naive Bayes

Text classification system to automatically classify notes and assignments using SVM and Naive Bayes with 87% accuracy.

statistical machine learning scikit-learn Support Vector Machine (SVM) Naive Bayes

GitHub Activity

My GitHub ↗ contributions over the past year. Colored squares represent days with commits.

GitHub contribution chart