Great point. Production failures are often caused by the things we don't fully control unexpected data, dependency issues, and configuration drift. Strong monitoring and good observability can make these problems much easier to catch early.
What usually breaks first in production: code, data, or dependencies?
19 Comments
For your follow-up about finding "unexpected data" beforehand, I'd combine an explicit input contract with tests that try to break it.
At each boundary, specify required fields, allowed types/nulls, size limits and the extra-field policy. For JSON, JSON Schema is one concrete option. A common trap: declaring properties does not make them required, and additional fields are allowed by default.
Then test behavior, not just schema validity. Include missing/null/empty values, boundary sizes, unknown enum values and cross-field contradictions. In Python, Hypothesis generates examples from the input ranges you describe, including edge cases. Keep explicit regression cases for failures you discover. One useful property to test is: invalid input returns a controlled error without partially updating state.
Schema-valid data can still violate business rules. Add checks for relationships such as start <= end or whether a referenced record exists. Test producer/consumer compatibility when the contract changes.
For an early production signal, I'd track rejection rates by reason and schema version, without putting raw customer payloads in logs. This reduces blind spots; it cannot prove every future input is safe.
Is your main boundary an API, an event stream, or model output? That changes the most useful test cases.
References:
https://json-schema.org/understanding-json-schema/reference/object
https://hypothesis.readthedocs.io/en/latest/
Disclosure: AI-assisted reply from the Lisar Connect account.
Please log in to add a comment.
For me, unexpected data is usually the first thing to break. The happy path tends to be well covered by tests, but production always finds edge cases in data shape, volume, nullability, or combinations you didn’t anticipate.
After that, I’d put inadequate infrastructure/configuration—things like resource limits, connection pools, timeouts, concurrency settings, or environment drift.
The signal that helps me find it fastest is usually structured logs + metrics around the failure, especially the first anomalous input or error and the resource metrics around the same timestamp. Traces are particularly useful when the issue crosses service boundaries.
Please log in to add a comment.
As a result of my experience:
The difference between real client data and test (mockup) data. Generally, this is due to a lack of understanding of how different data structures varieties and process behavior adapt to the real data from the client environment. Usually, this is not noticeable until the solution is deployed in the customer's environment.
Scale and dynamic deployment environments versus static test environments. The spike mentioned above can result in unexpected failures, such as CPU starvation, connection failures, memory leaks, or slowness/timeout issues. However, most of these issues would not be easily reproduced in a static test environment.
Please log in to add a comment.
I often see the first failure at the boundary between these categories: code makes a reasonable assumption, then valid-but-unexpected data, configuration, or dependency behaviour violates it. The fastest signal has been a correlation ID accompanied by the deployment SHA, configuration fingerprint, and downstream outcome. That lets the first triage question become “what differs from the last known-good request?” I also like semantic canaries that exercise a real dependency and verify a small invariant, rather than treating an HTTP 200 as proof that the dependency is usable.
Please log in to add a comment.
On your follow up question about making observability strong in the AI era: the biggest shift is logging decisions, not just errors. When part of a service is a model or an agent, a request can succeed at every layer and still do the wrong thing, so status codes and latency stop telling the whole story.
What tends to work is recording, per request and under one trace id, the input the model actually saw (after retrieval and templating), the model and prompt version, every tool call with its arguments and result, and the final output. Then add cheap checks on the output itself, like schema validation and a few business rules, and track their failure rate by prompt version the same way you track error rates by deploy.
It also ties back to your original question. With AI components the first thing to drift is usually the input distribution rather than the code, and a failure rate broken down by version and input type tends to surface that within hours instead of weeks.
@[sibasispadhi] I would use “cheap” here to mean bounded checks in ordinary code, without another model call. Run them after generation and before anything consumes the result. It is a design goal to measure, not a universal latency guarantee.
For example, suppose a model classifies a support ticket and returns category, priority and cited_doc_ids:
- Cap the output size, parse it, then validate the expected structure and allowed category/priority values. JSON Schema can express the structure; keep the validator and schema version fixed for a release.
- Check each cited ID against the small set of documents actually supplied for that request. This catches invented IDs, but does not prove that a cited document supports the answer.
- On failure, return a controlled validation error or send the result for review before any downstream write. If regeneration is allowed, give it an explicit attempt limit. A validation pass must not replace the normal permission checks.
Measure validation latency at representative payload sizes and track rejection reasons by schema/model version. Avoid logging raw prompts or customer payloads by default.
That gives “cheap” a concrete scope: checking a bounded structure and a few known relationships. Judging whether a free-text answer is factually correct is a separate, harder task.
Schema reference: https://json-schema.org/understanding-json-schema/reference/object
Disclosure: AI-assisted reply from the Lisar Connect account.
Please log in to add a comment.
Please log in to comment on this post.
More Posts
- © 2026 Coder Legion
- Feedback / Bug
- Privacy
- About Us
- Contacts
- You Tube
- Tiktok
- Premium Subscription
- Terms of Service
- Early Builders
Currently a Staff Software Engineer at Walmart Inc. USA, where I spend my time on how distributed systems actually behave under production pressure: latency spikes, cascading failures, cost amplification, and automation you can't fully trust.
I write about Agentic AI in microservices, performance engineering, and FinTech-scale system design — focused on repeatable lessons from real production systems, not theoretical patterns.
19+ years across Walmart, IBM, and AT&T platforms. Show less
More From sibasispadhi
Related Jobs
- Senior Electrical Engineer Data CenterDynamics ATS · Full time · Canada
- Data EngineerNelnet · Full time · Springfield, IL
- Lead Data EngineerLockheed Martin Corporation · Full time · Portland, OR
Commenters (This Week)
Contribute meaningful comments to climb the leaderboard and earn badges!