Summary (สรุป)
เมื่อซอฟต์แวร์ของเรากลายเป็นสิ่งสำคัญมากขึ้นต่อชีวิตของผู้ใช้ แรงขับเคลื่อนในการพัฒนา resiliency ของซอฟต์แวร์ที่เราสร้างก็เพิ่มขึ้นตาม แต่ตามที่เราได้เห็นในบทนี้ เราไม่สามารถบรรลุ resiliency ได้ เพียงแค่ ด้วยการคิดถึงซอฟต์แวร์และ infrastructure ของเรา เราต้องคิดถึงผู้คน กระบวนการ และองค์กรของเราด้วย ในบทนี้เราได้ดูสี่แนวคิดหลักของ resiliency ตามที่ David Woods ได้อธิบายไว้:
Robustness
ความสามารถในการรองรับความปั่นป่วนที่คาดการณ์ไว้ล่วงหน้า
Rebound
ความสามารถในการฟื้นตัวหลังเหตุการณ์ร้ายแรง
Graceful extensibility
เรารับมือกับสถานการณ์ที่ไม่คาดคิดได้ดีแค่ไหน
Sustained adaptability
ความสามารถในการปรับตัวอย่างต่อเนื่องต่อสภาพแวดล้อม ผู้มีส่วนได้ส่วนเสีย และความต้องการที่เปลี่ยนแปลงไป
เมื่อมองแคบลงมาที่ microservice พวกมันให้เราหลายวิธีในการพัฒนา robustness ของระบบเรา แต่ความทนทานที่พัฒนาขึ้นนี้ไม่ได้มาฟรีๆ—คุณยังต้องตัดสินใจว่าจะใช้ตัวเลือกไหน stability pattern สำคัญๆ อย่าง circuit breaker, time-out, redundancy, isolation, idempotency และอื่นๆ ล้วนเป็นเครื่องมือที่คุณมีไว้ใช้ แต่คุณต้องตัดสินใจว่าจะใช้เมื่อไรและที่ไหน นอกเหนือจากแนวคิดแคบๆ เหล่านี้ เราก็ยังต้องคอยระวังสิ่งที่เราไม่รู้อยู่ตลอดเวลาด้วย
คุณยังต้องหาว่าคุณต้องการ resiliency มากแค่ไหน—และนี่มักจะถูกกำหนดโดยผู้ใช้และเจ้าของธุรกิจของระบบคุณเสมอ ในฐานะนักเทคโนโลยี คุณสามารถเป็นผู้รับผิดชอบวิธีการทำสิ่งต่างๆ ได้ แต่การรู้ว่า resiliency แบบไหนที่จำเป็นต้องอาศัยการสื่อสารที่ดีและใกล้ชิดบ่อยครั้งกับผู้ใช้และเจ้าของผลิตภัณฑ์
กลับมาที่คำพูดของ David Woods จากก่อนหน้านี้ในบทนี้ ที่เราใช้ตอนพูดถึง sustained adaptability:
ไม่ว่าเราจะทำได้ดีแค่ไหนในอดีต ไม่ว่าเราจะประสบความสำเร็จมากแค่ไหน อนาคตก็อาจแตกต่างออกไป และเราอาจไม่ได้ปรับตัวให้เข้ากับมันดีพอ เราอาจเปราะบางและง่อนแง่นเมื่อเผชิญกับอนาคตใหม่นั้น
การถามคำถามเดิมซ้ำแล้วซ้ำเล่าไม่ได้ช่วยให้คุณเข้าใจว่าคุณพร้อมสำหรับอนาคตที่ไม่แน่นอนหรือไม่ คุณไม่รู้ในสิ่งที่คุณไม่รู้—การใช้แนวทางที่คุณเรียนรู้และตั้งคำถามอย่างต่อเนื่องเป็นกุญแจสำคัญในการสร้าง resiliency
Stability pattern แบบหนึ่งที่เราดูไป คือ redundancy สามารถมีประสิทธิภาพมาก ไอเดียนี้เชื่อมโยงได้ดีกับบทถัดไปของเรา ที่เราจะดูวิธีต่างๆ ในการ scale microservice ของเรา ซึ่งนอกจากจะช่วยเรารับมือกับโหลดที่มากขึ้นแล้ว ยังเป็นวิธีที่มีประสิทธิภาพในการช่วยเรา implement redundancy ในระบบของเรา และด้วยเหตุนี้จึงพัฒนา robustness ของระบบของเราด้วย
1 David D. Woods, "Four Concepts for Resilience and the Implications for the Future of Resilience Engineering," Reliability Engineering & System Safety 141 (September 2015): 5–9, doi.org/10.1016/j.ress.2015.03.018.
2 Jason Bloomberg, "Innovation: The Flip Side of Resilience," Forbes , September 23, 2014, https://oreil.ly/avSmU .
3 See "Google Uncloaks Once-Secret Server" by Stephen Shankland for more information on this, including an interesting overview of why Google thinks this approach can be superior to traditional UPS systems.
4 Michael T. Nygard, Release It! Design and Deploy Production-Ready Software , 2nd ed. (Raleigh: Pragmatic Bookshelf, 2018).
5 The terminology of an "open" breaker, meaning requests can't flow, can be confusing, but it comes from electrical circuits. When the breaker is "open," the circuit is broken and current cannot flow. Closing the breaker allows the circuit to be completed and current to flow once again.
6 Kripa Krishnan, "Weathering the Unexpected," acmqueue 10, no. 9 (2012), https://oreil.ly/BCSQ7 .
7 Russ Miles, Learning Chaos Engineering (Sebastopol: O'Reilly, 2019).
8 Kate Aubusson and Tim Biggs, "Major Telstra Mobile Outage Hits Nationwide, with Calls and Data Affected," Sydney Morning Herald , February 9, 2016, https://oreil.ly/4cBcy .
9 I wrote more about the Telstra incident at the time—"Telstra, Human Error and Blame Culture," https://oreil.ly/OXgUQ . When I wrote that blog post, I neglected to realize that the company I worked for had Telstra as a client; my then-employer actually handled the situation extremely well, even though it did make for an uncomfortable few hours when some of my comments got picked up by the national press.
10 John Allspaw, "Blameless Post-Mortems and a Just Culture," Code as Craft (blog), Etsy, May 22, 2012, https://oreil.ly/7LzmL .