Full-stack developer with an MCA and 3+ years of dedicated Python development, now focused on large-language-model evaluation — designing tasks and tests that reveal where AI coding agents fail. I build in Python, .NET, and SQL, and I write about it.
I’m a software developer and AI-data specialist with a Master of Computer Applications and a long track record across software engineering, database work, and quality assurance. I’ve built full-stack web applications and backend systems in Python, .NET, C#, and Java, and tuned complex SQL and Oracle PL/SQL workloads for performance.
More recently my focus has shifted to AI evaluation: auditing large-language-model output for factual and technical correctness, and authoring benchmark tasks that genuinely challenge frontier coding models — designing the task, defining what “solved” means, and writing verifiers that accept every valid solution while rejecting subtly broken ones. It sits right where software engineering meets AI quality, and it’s the work I find most interesting.
A dependency-free harness for scoring LLM answers against a gold set, with per-question match rules (exact / contains / numeric), gold-file validation, actionable failure output, and a CI-friendly exit code.
A compact task-tracker Web API in a single file — full CRUD with validation and correct status codes — plus integration tests that boot the app in memory and exercise the real endpoints.
Runnable examples that turn a slow row-by-row load into fast batched processing with BULK COLLECT and FORALL, plus SAVE EXCEPTIONS error handling — demonstrated on a 200k-row table.
A full-stack web portal, built during earlier software-development work, for submitting municipal complaints and tracking their resolution — with an admin dashboard that routes each complaint to the right department and monitors real-time status. Full-stack build across backend, database, and UI.
More at dev.to/zahid23saim
Languages: English, Hindi, Tamil, Marathi — professional fluency (native / near-native).
Open to remote, project-based software and AI-evaluation work.