Posts by Ayush_SIngh

@Ayush_SIngh

Ayush Singh

Building open source infrastructure to detect LLM hallucinations, prompt attacks...
India github.com/AyushSingh110 Joined May 2026
2k Points57 Badges16 Connections16 Followers16 Following

Comments by Ayush_SIngh

3 hours Articles 3 min read
OpenAI just released GPT-6 Astra. But one number immediately stands out: 99.9% on ARC-AGI-3. Considering that ARC benchmarks were created specifically to study the gap between current AI systems and general intelligence, that number sounds almost a...
Aug 27 Articles 4 min read
Earlier this year, Anthropic gave Apple, Microsoft, Google, Amazon and a handful of other companies access to an AI agent called Mythos and turned it loose on their code. Project Glasswing, as the initiative became known, found bugs that had survived...
post-cover-25476
Jul 31 Articles 5 min read
If you have shipped an LLM agent, you have watched one go sideways in real time. Halfway through a task, it takes one bad step, and now it is confidently marching toward a wrong answer. The whole industry has gotten good at detecting that moment sp...
Jun 16 Articles 5 min read
AI agents rarely fail in a clean, obvious way. They do not always crash. They do not always throw an error. They do not always say, "I could not complete the task." Sometimes they fail more quietly. They give a confident answer with weak evidence....
post-cover-20606
Jun 12 Articles 3 min read
Nobody tells you what building alone actually feels like. The blog posts make it sound clean. You have an idea, you build it, you launch. Maybe you hit some technical walls, you push through, and eventually things work out. What they skip is the pa...
post-cover-20218
Jun 12 Articles 3 min read
Nobody tells you what building alone actually feels like. The blog posts make it sound clean. You have an idea, you build it, you launch. Maybe you hit some technical walls, you push through, and eventually things work out. What they skip is the pa...
post-cover-20218
Jun 6 Articles 1 min read
Most security systems are evaluated on attacks they have already seen. I decided to test mine on ones it hadn't. The Setup I built FIE "an open-source adversarial prompt detector for LLMs". 11 detection layers run in parallel on every incoming prompt...
Jun 4 Articles 3 min read
Most developers know they shouldn't use production data in non-production environments. But knowing and doing are two different things — and the gap between them is getting more expensive. According to Nick Mathison, Senior Product Manager for Delph...
post-cover-19698
May 15 Articles 2 min read
Most Dev's Journey stories start with a first line of code. Mine starts somewhere different. I came from marketing. And somewhere along the way, I became a technology writer — not because I could build software, but because I was genuinely curious a...
post-cover-17847
May 14 Articles 2 min read
Most Dev's Journey stories start with a first line of code. Mine starts somewhere different. I came from marketing. And somewhere along the way, I became a technology writer — not because I could build software, but because I was genuinely curious a...
post-cover-17847
chevron_left

Latest Jobs

View all jobs →