News: 1713326385

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Tencent Cloud to revisit design after circular dependencies slowed emergency API fix

(2024/04/17)


Tencent Cloud has apologized for an outage that impacted customers last week – an unusual act by a Chinese cloud – and signalled it will review some aspects of its ops in the hope of avoiding future incidents of this nature.

A [1]WeChat post from China's number three cloud – Alibaba Cloud leads the market, ahead of Huawei, with Baidu in fourth place – revealed that on April 8 it updated configuration data for one of its APIs.

The change was a dud, and the API became unavailable. Some Tencent Cloud platform-as-a-service offerings services that rely on it therefore became unreliable, effectively cutting them off from the Tencent Cloud.

[2]

The cloud provider was able to fix the mess in just 87 minutes, but has apologized to the 1,957 customers who reported failures.

[3]

[4]

The triage post explains that the failure was caused by an update to the API that didn't consider compatibility.

Changes to the API's interface protocol meant that apps targeting the old version produced nonsense data that spread across Tencent Cloud and meant the API became unstable.

[5]

Tencent Cloud would usually roll back this sort of change. But to do so, it needed the very same API it had just broken.

[6]Tencent explores a future where HPC, quantum, cloud and edge have converged

[7]Tencent, Meta, alliance reportedly strains over differing VR visions

[8]Baidu joins Tencent in downplaying impact of US chip bans on AI ambitions

[9]Alibaba Cloud reveals network telemetry tool that helped cut number of engineers needed by 86%

The cloudy concern admitted that it just didn't test this release properly – it ignored some of its own version change processes, didn't conduct proper sandbox tests, and now realizes its change management processes probably need some work.

The outfit has pledged to redesign bits of its cloud to detect abnormal changes and terminate them before they spread, conduct drills to improve its incident response, and offer alternative APIs should the interfaces fail.

Tencent Cloud is not alone in breaking its own cloud with an update. The Register has reported on similar messes at [10]Google , [11]AWS , and [12]Microsoft .

One important difference, however, is that Tencent Cloud is reputedly so intolerant of outages that it has been known to [13]fire staff responsible for resilience after incidents.

[14]

Chinese media [15]suggest that approach may have backfired – sources allege job cuts at Tencent Cloud contributed to this incident. ®

Get our [16]Tech Resources



[1] https://mp.weixin.qq.com/s/2e2ovuwDrmwlu-vW0cKqcA

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Zh@dvyI47O4KquZoqiI2hQAAAME&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zh@dvyI47O4KquZoqiI2hQAAAME&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Zh@dvyI47O4KquZoqiI2hQAAAME&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Zh@dvyI47O4KquZoqiI2hQAAAME&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[6] https://www.theregister.com/2024/01/29/tencent_huawei_tech_predictions/

[7] https://www.theregister.com/2024/01/22/apac_tech_news_roundup/

[8] https://www.theregister.com/2023/11/22/baidu_q3_2023/

[9] https://www.theregister.com/2024/04/16/alibaba_cloud_zoonet_network_telemetry/

[10] https://www.theregister.com/2020/08/25/gmail_outage_root_cause/

[11] https://www.theregister.com/2020/11/30/aws_outage_explanation/

[12] https://www.theregister.com/2023/09/04/microsoft_australia_outage_incident_report/

[13] https://www.theregister.com/2023/04/14/tencent_outage_firings/

[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_offprem/front&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Zh@dvyI47O4KquZoqiI2hQAAAME&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[15] https://www.caixinglobal.com/2024-04-10/job-cuts-might-be-behind-outage-at-tencent-cloud-source-says-102184587.html

[16] https://whitepapers.theregister.com/



"There's a hole in my bucket."*

Bebu

《Tencent Cloud would usually roll back this sort of change. But to do so, it needed the very same API it had just broken.》

Have to wonder how much of any organisation's services and infrastructure has a cyclic graph of dependencies. Trees are good but dags (directed acyclic graphs) will do. ;)

Possibly not a triumph of devops but perhaps not. :)

* [1]a childhood memory.

[1] https://www.songsforteaching.com/folk/theresaholeinthebucket.php

Roll back failed as API unavailable

Anonymous Coward

Just allow me a moment of schadenfreude - HAAAAAR!

I own seven-eighths of all the artists in downtown Burbank!