Vendia | 10 ways to make your software better when building for cloud scale
10 ways to make your software better when building for cloud scale
- November 17, 2022
Posted by Nitesh Arora
Scale. It’s one of the most common stressors to complex software systems. So how do public cloud services companies do it? And what differentiating practices can we apply from their toolkit, even when developing solutions that don’t require “cloud scale”?
Before exploring the practices needed for cloud scale, it’s important to set a baseline for any high-performing software team. The practices outlined in the DevOps Research Assessment (DORA) are a great starting point. The technical practices, Lean product development, Lean management, and culture and work environment approaches in DORA are applicable no matter the intended scale of your software system.
But there also must be something new or different to bolster those practices to achieve cloud scale. The 10 practices I’ve observed from the Vendia engineering team — those that seem to differentiate regular software from cloud-scale software — are grouped into four categories:
- Knowledge – What you and your team members know
- Operations – How you and your team members handle operations
- Development – How you and your team members build software
- Helping others – What you and your team members do to share all of the above with others
Together these categories, and the practices within them, allow a software team to increase the speed with which they can develop new features, increase the quality of the software they deliver and increase their likelihood of achieving the scale they desire.
Knowledge
Practice No.1 – Hire or develop distributed systems expertise
- In regular development teams, distributed systems expertise is often concentrated, which limits the efficiency and effectiveness of the team.
- In cloud-scale teams, distributed systems expertise is pervasive.
Software teams building for cloud scale must deeply understand the common pitfalls of distributed systems and their solutions. While reading about them is a start, experience dealing with the failures instills a level of understanding that is unmatched. Likewise, calling upon experiences, both successful and unsuccessful, gives teams the resources to find ingenuitive ways to solve those challenges, even at a moment’s notice.
Practice No. 2 – Get very familiar with the libraries and cloud services you use
- In regular development teams, library and cloud service selection is based on what’s popular and what works most quickly for simple cases. This may result in selections that yield ugly failure scenarios.
- In cloud-scale teams, cloud service and library selections involve deep exploration of source code (when available), documentation, and a large focus on experimentation, benchmarking, and prototyping (including failure modes).
You may have heard the adage, “An amateur practices until they can do a thing right; a professional, until they can’t do it wrong.” Reframed for the software context, the quote looks something like this: “An amateur integrates a library until they get it to work; a professional, until it will always work.”
Spending an appropriate amount of time to select the right library, one that works but also fails gracefully, is a great realization of DORA’s team experimentation practice. The same selection rigor applies to public cloud services as well. Their inherent limits and failure modes must be fully understood to properly design and build for those potential failures.
Operations
DORA covers operations extensively, including its measurement capabilities. What differs about cloud-scale teams in the area of operations are not the practices themselves, but the level of sophistication and completeness applied.
Practice No. 3 – Provide deep and complete operational visibility
- In regular development teams, operational visibility is not given the same time and attention as feature development, which shortchanges instrumentation and often relies on insufficient tooling for metric correlation across complex distributed systems.
- Cloud-scale teams understand that operational excellence requires deep and complete operational visibility. They know that the absence of evidence is not the same thing as the evidence of absence.
Monitoring and observability are already well-established imperatives. The guidance in DORA outlines several areas of focus, including instrumentation and correlation. Instrumentation means that code must be added to a system to expose its inner state. Correlation means that metrics collected from the application and its underlying systems can be used to explain behavior.
Practice No. 4 – Create robust operational automation and tooling
- Regular teams use automation where they know others use automation. They spend time and attention on internal tools, but they don’t apply the same design and code quality rigor they would for new feature development.
- Cloud-scale teams understand that automation should be applied anywhere and everywhere repeatable processes exist, including to software development.
The purpose behind automation is to reduce toil, increase the likelihood of success and make important but painful processes simple, reliable, and, as a result, more frequently executed.
Practice No. 5 – Hold rigorous on-call reviews
- Regular teams limit performance reflection to retrospectives, which can become rote practice if action items are not captured and closed quickly.
- Cloud-scale teams know that operational visibility provides a wealth of data that can be gleaned to improve design, implementation and team performance, provided sufficient time and energy is provided to inspect and learn from that data.
Practice No. 6 – Always be canary-ing (ABCs)
- Regular development teams may opt for a blue/green deployment strategy since it is easier to mechanize than canary releasing.
- Cloud-scale teams work in small batches and leverage leading operational practices.
Development
There are endless resources available to help you and your team improve software development practices.
Practice No. 7 – Engineer specifically for failure scenarios
- Regular teams develop code that catches and handles basic failures using basic techniques, possibly overlooking the less common but more impactful failure scenarios.
- Cloud-scale teams spend the majority of their design and development on avoiding, handling, testing and automatically recovering from failure scenarios.
Practice No. 8 – Apply frequent and extensive code reviews
- Regular development teams use code reviews to improve code quality.
- Cloud-scale teams use code reviews to improve code quality, software design, operational visibility, user experience, automation, and tooling.
Helping others
Practice No. 9 – Provide developer environments on Day 1
- Regular teams are often stifled by the environments in which they’re allowed to operate, and new team members often need to wait days or weeks for access to the systems needed to make them productive.
- When you join a cloud-scale team, you have the tools you need to ramp up quickly, deliver value frequently and experiment freely.
Practice No. 10 – Give back through open source
- Regular teams heavily rely on open source libraries but rarely contribute back.
- Cloud-scale teams use open source libraries extensively and create their own, knowing that, by doing so, they are both improving their craft and the craft of others.
Key takeaways
Cloud-scale development teams are similar to regular development teams in many ways. In some cases, cloud-scale teams apply an additional level (or two) of rigor to existing practices. In others, they expand the scope of a common practice or apply it to areas not previously intended. When combined with an excellent set of foundational practices, the 10 practices defined above are enough to differentiate cloud-scale teams from all the others.