Skip to content
View WatchTree-19's full-sized avatar

Block or report WatchTree-19

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
WatchTree-19/README.md

Hi, I'm Sandeep

This is my open research page, feel free to reach out by email, if you feel I could help out! (Columbia University, alumnus), UK/US-based.

I place myself as an Methodological researcher.

[ What I'm building

  • Independent writing on AI evaluation methodology, observability, and the structural overlap between quant trading and LLM eval.
  • Asymmetric-information solutions in ML evaluation, surfacing what labs know internally about benchmark noise and drift.
  • Calibration tooling for benchmark drift, distinguishing genuine model improvement from eval movement.

[ Currently working on

  • A foundational essay on production observability for LLM agents.
  • A weekly paper digest series on alignment, evaluation methodology, and AI safety research.
  • "Benchmark crowding": mapping factor decay in quant finance to benchmark saturation in LLM evaluation.

[ Around the web

Pinned Loading

  1. llm-judge-calibration llm-judge-calibration Public

    Measure how much your LLM judges actually agree. Inter-judge agreement metrics for LLM-as-a-judge evaluations.

    Python

  2. UKGovernmentBEIS/inspect_ai UKGovernmentBEIS/inspect_ai Public

    Inspect: A framework for large language model evaluations

    Python 2.4k 626

  3. EleutherAI/lm-evaluation-harness EleutherAI/lm-evaluation-harness Public

    A framework for few-shot evaluation of language models.

    Python 13.5k 3.5k

  4. pola-rs/polars pola-rs/polars Public

    Extremely fast Query Engine for DataFrames, written in Rust

    Rust 39.2k 3k

  5. pixie-io/pixie pixie-io/pixie Public

    Instant Kubernetes-Native Application Observability

    C++ 6.5k 498

  6. NVIDIA/garak NVIDIA/garak Public

    the LLM vulnerability scanner

    Python 8.6k 1.1k