Skip to main content
Cloud

First jobs: How a Visa engineer learned to stay calm on call

Sai Joshitha Kathari recalls the importance of “systematic” thinking.

4 min read

TOPICS: Cloud / IT Infrastructure & Operations / Kubernetes (Cloud)

It was after hours, and Sai Joshitha Kathari was on call—and on her own.

Kathari, a then-new hire at digital transformation company Atos Syntel, received an alert that would likely frighten any IT newbie: a production service was acting up. Several Kubernetes application pods were repeatedly failing, reducing the service’s capacity and requiring immediate troubleshooting to prevent a broader outage.

Kathari had just finished school. Her responsibilities, as a new member of Atos’s DevOps engineering team, included supporting production applications running on Red Hat OpenShift, an enterprise Kubernetes platform. Because customers required those enterprise services, Kathari and colleagues had to be on call at times, including working the night shift to handle any potential infrastructure or application problems.

Kathari recalled her first alert while speaking to IT Brew.

“Honestly, I was nervous. I had just graduated, and suddenly I was responsible for helping support systems that real people depend on. That’s a very different feeling from working on assignments in college,” Kathari told us. “I remember thinking, ‘What if I don’t know the answer? What if I make the wrong decision?’”

Kathari did make the correct decisions, solving the problem well before the morning IT crew arrived. Now a senior site reliability engineer at Visa, almost 20 years later, she still considers that early emergency a significant moment in her career.

“The biggest thing I learned wasn’t a technical skill at all. It was how to stay calm,” Kathari said.

Step by step. After receiving the alert and taking a deep breath, Kathari walked herself through important remediation steps, starting with logging in to review the pod status, along with monitoring data and application logs.

The immediate issue appeared to be resource pressure affecting the pods’ stability, she recalls, meaning that the pods did not have sufficient computing resources available to operate reliably.

Top insights for IT pros

From cybersecurity and big data to cloud computing, IT Brew covers the latest trends shaping business tech in our 4x weekly newsletter, virtual events with industry experts, and digital guides.

By subscribing, you accept our Terms & Privacy Policy.

Kathari restarted the affected pods, adjusted their resource allocation (increasing CPU and memory), and monitored the application until the pods were stable again.

That may sound straightforward, but an on-call emergency is no easy A.

“In college, you are basically solving problems where you have time to think,” she said. “In production, it’s different. There is no answer key, and there may be customers actively affected while you are investigating.”

The help desk is often considered a valuable first stop in an IT pro’s career path, thanks largely to its exposure to complicated, breakable technologies, supportive peers, and opportunities to diagnose problems and troubleshoot. An on-call service tech also has a chance to handle crises on their own.

When recalling the late-night alert, Kathari remembers shifting from fear to focus. “Instead of worrying about what could go wrong, I started thinking, ‘What’s the first thing I need to check?’” she said, which meant reviewing dashboards and logs.

That step-by-step thinking comes in handy for Kathari today, where she’s tasked with automating and validating production services, as well as resolving any disruptions. Similar steps apply: trust the data, communicate clearly, and solve the problem. When validating production changes or investigating a disruption today, she narrows the scope: Is the problem affecting one application instance, one service, one cluster, or a larger dependency? Then, she reviews monitoring signals, logs, deployment changes, and application behavior.

“Engineering isn’t about knowing everything; it’s about knowing how to approach a problem systematically,” Kathari said. “Even today, as a SRE, that’s probably the most important skill I use every day.”

About the author

Billy Hurley

Billy Hurley has been a reporter with IT Brew since 2022. He writes stories about cybersecurity threats, AI developments, and IT strategies.

Top insights for IT pros

From cybersecurity and big data to cloud computing, IT Brew covers the latest trends shaping business tech in our 4x weekly newsletter, virtual events with industry experts, and digital guides.

By subscribing, you accept our Terms & Privacy Policy.