Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

(gist.github.com)

53 points | by wsxiaoys 1 hour ago

4 comments

  • wongarsu 49 minutes ago
    That writing style might be a tad too tense

    If I got it correct (appending B from https://stolen-thoughts.com/paper.pdf is essential) they are the authors of the well-known exploit to recover readable CoT from OpenAI and Anthropic models. They use that to find hints of distillation, by running a benchmark with a SotA model, recovering the CoT, then taking the first 1% of the CoT and running the open-source model as if that was the start of its own CoT. In the paper they found that Kimi-K3 gets a lot closer to Claude 4.8 answers when prefilled with the start of Claude 4.8 reasoning, suggesting that Claude 4.8 was used in its post-training. This blog post is the follow-up with results that suggest that Qwen3.8 was post-trained with the help of GPT-5.5 Pro (or some similarly responding GPT model, it's unclear how many models they tested)

  • 7734128 55 minutes ago
    The problem with this is obviously that the only GPT 5.5 thoughts that we have access to are from stolen thought.

    Qwen 3.8 0902 was trained after the release of the paper on August 10, so it should have seen those specific thoughts.

    • usernomdeguerre 32 minutes ago
      seems like only the companies in question could run this sort analysis long-term; since they have full access to their CoTs not in public datasets.
      • verdverm 3 minutes ago
        and we have to "trust them bro" to be fair and accurate, something I am very unlikely to do given their other false / misleading statements to date
    • refulgentis 31 minutes ago
      The thoughts trick was known before their paper / August.

      I "independently" "invented" it for the first Anthropic reasoning models because the API required you have thoughts for each assistant message. My app lets you switch AIs within a chat, and their API used to require thinking for all messages if thinking was enabled, so I needed to get a valid thinking stub to insert.

      Time has flew by for me the last 3 years, but, I'd guess it's been at least 18 months. And IMHO it wasn't very complicated to work through how to do once you were dead set on making it happen, this I expect it was well-known to distillers before the paper.

  • jari_mustonen 32 minutes ago
    > Qwen barely moved toward Opus 4.8 in the earlier experiment, but moved by +20.58 points toward GPT-5.5 Pro here, including a large effect on the private synthetic puzzles. The data suggest that Qwen may have learned from GPT-5.5 Pro, or from a closely related GPT model, rather than from Opus.

    How does this suggest anyting of the sorts?

  • CamperBob2 28 minutes ago
    News flash: people who scraped the Internet without permission to build their product complain when something vaguely similar is done to them. Water still wet, sky still blue. Film at 11.
    • verdverm 1 minute ago
      I would not be surprised in the slightest if we later find out they are running those same open models to find useful traces or bits to incorporate into their own training. Lots of rules for thee but not for me from Big Ai
    • noir_lord 10 minutes ago
      I'd send them the worlds smallest violin but Rufus is getting in the way of me finding it.