Snippet
Priyank Upadhyay is from Asansol Engineering College in West Bengal, not IIT or Stanford. He led mobile apps and backend at Yellow.ai, worked on Kubernetes multi-cluster infrastructure at Avesha, and spent years in open-source communities before starting a company. RubixKube is an AI-native Site Reliability Intelligence platform. It’s been quietly answering that question ever since.
When Priyank Upadhyay explained what he was building to the CTO of a leading AI company, the response was simple
“How can I log in?”
That conversation became RubixKube’s first real design partnership. The team deployed the prototype in a staging environment, and the feedback helped shape the product into what it is today. One early engagement revealed a hidden cloud cost issue that would typically have taken days of engineering investigation to identify. RubixKube found it in five minutes, and the planning and migration were completed over a weekend.
The same conversation, everywhere
Priyank didn’t come to this problem from a product spec. He came from talking to engineers at scale.
In 2003, Google introduced the idea of the site reliability engineer. The premise: if you let developers inside the server room, gave them the tools and authority to automate what had been done manually, you could build systems that rarely broke and recovered fast when they did. The split was specific. Thirty percent of an SRE’s time on incident management. Thirty percent on automation. Forty percent on new capability.
The industry adopted the title. It didn’t adopt the split.
At his previous experiences, in conversations with NVIDIA, Google, and Oracle’s engineering teams, Priyank kept finding the same thing. SREs spending 90% of their time firefighting. When something broke at 2am, someone got paged. They opened five tabs. They ran terminal commands, went to Slack to find the right person, pulled up Confluence docs that might be relevant, tried to reconstruct from fragmented signals why something that worked yesterday had stopped working tonight. The automation bucket, the 30% that was meant to make the whole system self-sustaining, never got funded. There was always a more urgent fire.
What stayed with him wasn’t just the inefficiency. It was what he calls amnesia. When a senior engineer or team lead leaves, they take with them a map of the system that nobody else fully holds. Which services depend on which. What config change three months ago might explain a failure today. What was discussed in a Slack thread two weeks back that never made it into the official documentation. A new hire starts from scratch. The postmortem from last quarter’s outage sits in a document nobody reads. Every major incident begins with a team that is collectively less informed about its own infrastructure than it should be.
Invisible software
RubixKube sits inside production infrastructure and builds that map. It connects to everything, PagerDuty alerts, Jira tickets, Confluence documents, Slack threads, GitHub, Grafana dashboards, the live services themselves. Everything it sees gets stored in what they call a Memory Engine: a connected record of how the infrastructure works, what has changed, and what has broken before. When something goes wrong, the answer is already in there.
Who owns this service? What was it doing last week? If it goes down tonight, what else breaks with it?
Priyank describes the product to customers as “invisible software.” You don’t have to use it. You plug it in, ten minutes in a standard Kubernetes environment, an hour for more complex stacks, and then forget it’s there. It comes to you when something is wrong: here is what broke, here is why, here is the specific action you need to take.

The trust model is staged by design. Nobody gives an AI system authority over production on day one. RubixKube starts in observation mode: recommendations only, every action requiring human confirmation. Once an engineering team has seen the recommendations be correct enough times, they start delegating the smaller, well-understood tasks, scale down this service, restart that pod. The autonomous loop expands from there, earned through track record rather than assumed.
The trust that accumulates is the moat. A Memory Engine that has watched a specific company’s infrastructure for eighteen months knows that company in a way no external tool can replicate. When a larger player adds “AI SRE” as a feature, it starts from nothing. When RubixKube is in year two at a customer, it has context that is structurally impossible to acquire from the outside.
Building the category before it has a name
Priyank doesn’t call what RubixKube does “AI SRE.” He coined a different term: Site Reliability Intelligence.
His reasoning is architectural. Calling it AI SRE presupposes the technology should behave the way an SRE behaves, same constraints, same workflow, same ceiling. That’s too small a frame. What he’s after is infrastructure that understands itself: systems that detect failure, explain it, and over time fix it without waiting for a human to be paged.
“If something else comes, AGI, or whatever follows LLMs, the goal doesn’t change,” he says. “The infra should be intelligent enough to know what is happening, why it’s failing, how to fix it, and possibly fix itself. A self-healing infra is the future I’m imagining.”
Earlier this year, they launched a Cursor and Claude Code plugin. The idea: if an engineer is writing code that will be deployed to production, they should be able to ask, from the IDE, while the code is still being written, what it’s going to break. The plugin connects to the live infrastructure context and flags the impact before the push. The Memory Engine that helps SREs diagnose failures after they happen, applied at the point when code is still editable.
Learning mindset over skillset
Priyank’s team runs across four countries. Most are in Bangalore. Two are in Nepal. One is in Dubai. One is in the US. The daily standup happens at 11pm Bangalore time because that’s when all the timezones overlap.
Hiring for a stack that requires AI, Kubernetes, and production SRE experience in the same person is one of the harder talent searches in Indian engineering. Priyank’s filter isn’t a skills checklist. He hires on a learning mindset first. Ownership second. And third, the one that sounds like a risk rather than a criterion until you’ve run an early-stage team, the mindset of wanting to own a startup someday.
He describes it as borrowed from his previous experience at Yellow.ai, where the culture explicitly blessed people for leaving to build their own thing. “They give you the approach of the corporate ladder is not the only direction you should be focusing on.” He brought that posture to RubixKube. The bet is that people who think like founders make a better team than people who don’t, and that the ownership that enables good work isn’t something you install with a culture document.
He’s building infrastructure that doesn’t forget. He’s hiring people who, if everything goes right, will eventually leave to build their own thing. He considers those the same idea.

