Cloudflare AKE cuts origin HelloRetryRequests from 52% to 3.7%

(blog.cloudflare.com)

46 points | by iamsyr 3 hours ago

5 comments

  • sandeepkd 1 hour ago
    TLDR;

    1. The TLS handshake involves a step to discover the commonly supported algorithms and can incur additional roundtrip if the first guess does not works out, its part of the protocol to keep it stateless

    2. Cloudflare is scanning all the origins on daily basis and storing the result for supported algorithms to save on the possible roundtrip time

    Whats missing in the article - They are saving on the *possible roundtrip latency, however they are not sharing the absolute lookup latency which now gets added to every connection

    • sophacles 1 hour ago
      What latency gets added? Presumably on cache miss they already need to look up "whats the origin for www.example.com" and get info about it. This is just a handful of bytes in that record.
  • greatgib 1 hour ago
    2 things comes to my mind reading this article:

    1) So they saved 15ms on the connection so that you can then wait 20s in their annoying nag screen before reaching the real website content.

    2) On the Monday they complain about the load on server by LLM scrapings compulsively your webserver an offer themselves as the internet guardian solution; and on Tuesday, they compulsively send useless requests to your servers so that they can save a few microseconds in the very first connection ever to your server. "For each TLS 1.3 capable origin, we run a series of a few lightweight TLS handshakes, each offering exactly one key agreement group: X25519, P-256, P-384, P-521, or X25519MLKEM768. [...] And because the active scanning happens outside your production traffic path, we confirm that both your origin and the network in between can handle connections with a stronger key agreement before any real traffic depends on it.

    • innocent_name 46 minutes ago
      >they compulsively send useless requests

      Those requests aren't useless; They clearly optimize. What a silly take.

      >to your servers

      To their customer's servers, right! Most people turn on other CF optimizations like h2/h3 to origin.

    • gonzalohm 59 minutes ago
      Agree on point 1. Now even tiny websites that wouldn't be a target for anyone are hiding behind Cloudfare. If you are so worried about people sending requests to your website then take it offline, that will get you 100% success preventing bots

      I'm so tired of having to do Captchas and waiting everywhere to access websites

      • Kodiack 52 minutes ago
        As someone who hosts a few “tiny websites” that are “hiding behind Cloudflare”, it’s because of some incredibly misbehaved botnet traffic that’s otherwise persistently scraping via residential proxies.

        This is a hobby for me. It’s for my enjoyment and it allows me to provide resources that others enjoy using. However, I’m not going to allow literally 99%+ of requests to be aggressive scrapers that won’t give up until you’ve got a heavy-handed solution in place.

        I hate it too, but the alternative is even more consolidation, so unfortunately this is just the reality right now and you’ll have to get over it until when/if things improve.

        • tredre3 18 minutes ago
          > I hate it too, but the alternative is even more consolidation, so unfortunately this is just the reality right now and you’ll have to get over it until when/if things improve.

          No, the actual alternative is learning to do it yourself. Unless you're a frequent target of UDP DDoS, there is almost nothing you can't defend against straight from the server itself.

          Spending time learning how to configure your firewall, set rate limits in your web server, tune your application for caching and further rate limits, and should all this fail install a self-hosted captcha/proof-of-work fence yourself, is definitely a fair amount of efforts when those things aren't part of the core hobby.

    • sophacles 1 hour ago
      2) Perhaps theres a slight difference between thousands of requests per day from untrusted entities and a single TLS handshake per day from a trusted one? (Note i understand you clearly don't trust cloudflare, but the people who sign up for cloudflare do trust them - that's the trust relationship i refer to here.)
  • chrismorgan 2 hours ago
    Genuine question: why wouldn’t they have been doing this already? It feels like obvious low-hanging fruit on a critical path, so I presume there’s something more to it than I’m imagining.
  • ttul 44 minutes ago
    Cloudflare is the canonical example of why you can't vibe-code infrastructure. Knowing that this optimization was even necessary, let-alone having the ability to build it, is something that doesn't become apparent until you're operating at considerable scale. Once there, of course if you're Cloudflare, you use coding agents to build it. But outside of these temples of scale, good luck even knowing it was needed.

    If you work at a SaaS of any kind, I think it's worthwhile considering what things will look like when scale is the only thing that is really defensible anymore.

    • asdfman123 32 minutes ago
      The irony is that everything about this speaks to vibe coding. The writeup was written by Claude, in a good way (E.g "while the milliseconds are important, that second part may matter more").

      I think things like this are now possible through vibecoding.

      • binsquare 19 minutes ago
        I like to think that the coding agents has basically raised the bar.

        Those who are above the bar can steer and add their expertise to hit a new level.

    • epistasis 15 minutes ago
      I'm very shocked at the very very bad design decisions with low performance data modeling that I see in genomics across all frontier models. It's stuff that even a new trainee typically wouldn't do, and the models are confident that they don't even present these key decisions as a choice that was made in their implementation plan. I've had to go through many many turns with Claude Code to convince that it made very stupid choices and that there are far more obvious and performant data models than shave off an order of magnitude on both data size and compute time.
    • doctorpangloss 38 minutes ago
      Brother, if you think Cloudflare isn't vibe coding features...
  • lucaprata 2 hours ago
    We had a probe that checked whether a connection was using post-quantum cryptography. It looked for "Cipher is" in the output. But when the handshake failed, it would print "Cipher is (NONE)." So it reported that github.com and amazon.com were compliant. For weeks.

    The success criterion was embedded in the errorline.