Dipankar Sarkar · Applied AI research & technical evaluation
Resolve the claim before committing to the system.
Applied research, reproducible evaluations and technical validation. Bring a precise question, a consequential decision and the evidence that needs to be tested.
Commission a study
When a decision depends on a technical claim, test the claim first. Each study has a protocol agreed before any results exist, a fee that does not depend on the outcome, and raw results you can inspect.
Commissioned study
AI Benchmark Replication Studies
Re-run a published benchmark or paper result under your workload, with raw results and a deviation report.
Commissioned study
AI Benchmark and Evaluation Study Design
Build a pre-specified test for a decision existing benchmarks do not measure, with baselines and failure conditions.
Commissioned study
Independent Model and Retrieval Comparison Studies
Run shortlisted models or retrieval methods on one workload; report uncertainty, input-regime effects and costs.
Commissioned study
Generated GPU Kernel Correctness Studies
Test generated GPU kernels with operator-aware oracles, calibrated tolerances and adversarial inputs.
Commissioned study
Empirical AI Feasibility Studies
A bounded experiment with stop criteria fixed in advance, ending in a proceed-or-stop recommendation.
Collaborate
Sponsored research and collaboration
Fund a defined research question, collaborate on a shared one, or invite a methods seminar.
Current research questions
All research →Research
Agent Evaluation and Coordination Research
Escalation, refusal to fabricate, multi-agent coordination, judge robustness and ranking stability, with papers, artefacts and limits.
Research
Generated-Code and GPU-Kernel Correctness Research
Correctness oracles, test-input generation, calibrated tolerances and a 26-op failure corpus for LLM-generated kernels.
Research
Federated Learning and Privacy Research
Federated learning under imbalanced data, and the evidence a distributed-ML privacy claim needs: baselines, metrics, threat models.
Research
Decentralised Systems and Protocol Research
DePIN incentives, composability and MEV: protocol research with stated assumptions, models, experiments and limits.
Methods you can reuse
Worked protocols with their limits stated, plus two free tools.
Method
Benchmark Leakage and Contamination Checks
Method
Uncertainty, Repeated Trials and Evaluation Claims
Method
Agent Failure Taxonomies and Incident Coding
Method
Evaluation Dataset Design and Provenance
Method
Pre-Specified Evaluation Plans
Method
Research Reproducibility Packages
Free tool
Evaluation Brief Builder
Free tool
Reproducibility Readiness Checklist
Featured Publications
View all →arXiv
How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
Dipankar Sarkar
SeT-LLM @ KDD 2026
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
Dipankar Sarkar
arXiv
The Correctness Illusion in LLM-Generated GPU Kernels
Dipankar Sarkar
arXiv
Before the Pull Request: Mining Multi-Agent Coordination
Dipankar Sarkar
VecDB@VLDB 2026 submission
An Input-Regime Audit of Conflict Detection for Retrieval-Augmented Generation
Dipankar Sarkar
arXiv
Navigating the Knowledge Sea: Planet-scale answer retrieval using LLMs
Dipankar Sarkar
arXiv
Viz: A QLoRA-based Copyright Marketplace for Legally Compliant Generative AI
Dipankar Sarkar
arXiv
Decentralized Deepfake Detection Blockchain Network using Dynamic Algorithm management
Dipankar Sarkar
arXiv
Generalised DePIN Protocol: A Framework for Decentralized Physical Infrastructure Networks
Dipankar Sarkar
FL-IJCAI'20
Fed-Focal Loss for imbalanced data classification in Federated Learning
Dipankar Sarkar , A Narang , S Rai
Packt Pub
Nginx 1 Web Server Implementation Cookbook — Packt Publishing (2011)
Dipankar Sarkar
Patent applications 25
Provisional applications filed with the Indian Patent Office — not granted patents
A Method and a System for Dynamically Modifying a Virtual World
A System to Generate and Dynamically Update a Virtual World Based on User Interest
A Method and System to Recommend a Navigation Path in a Virtual World
Method and System to Generate Animated Audio-Visual Content
System and Method for Partitioning a Neural Network Model for Offloading Computational Load
Method for Converting Static Graphic into Animated Graphic
Method for Dynamic Content Generation and Device Thereof
A Method and System for Partitioning a Social Network Group
Recent & Upcoming Talks
Decentralized AI: Privacy, Fairness, and the Future of Machine Learning
AI & Web3 Summit 2024
A keynote on the convergence of federated learning, blockchain, and decentralized systems for privacy-preserving AI
Building the Decentralized Data Economy: From Theory to Practice
Web3 Infrastructure Conference
How DePIN and blockchain technology are creating new models for data ownership and monetization
MEV and Fair Value Distribution in Ethereum
DeFi Security Summit
Technical workshop on Maximal Extractable Value (MEV) and designing fairer blockchain protocols
Fed-Focal Loss for Imbalanced Data Classification in Federated Learning
International Workshop on Federated Learning for User Privacy and Data Confidentiality (FL-IJCAI'20)
A presentation on applying Focal Loss to Federated Learning for handling imbalanced data classification
Recent Posts
View all →Fair Ordering: Taming MEV on Ethereum
Maximal Extractable Value comes from a validator freedom to order transactions within a block, and every mitigation approach constrains that freedom differently.
Why GPU Kernel Benchmarks Miss Real Bugs
Passing a correctness benchmark and being correct in production are different claims for LLM-generated GPU kernels — the gap comes down to which inputs get sampled.
Diagnosing Conflicts in RAG Pipelines
Why retrieval-augmented generation fails when context disagrees with parametric knowledge, and a checklist for diagnosing which failure mode you are looking at.
From AutoPrompt to TextGrad: A Prompt Survey
A chronological survey of automated prompt optimization 2020–2025: AutoPrompt, APE, OPRO, EvoPrompt, DSPy, TextGrad, PromptAgent, and how to choose between them.