4 comments

  • astrobiased 1 hour ago
    Is this in any way similar to Goodfire's work? https://www.goodfire.ai/research/rlfr#
    • HenryNdubuaku 1 hour ago
      Thats an interesting outlook, loosely similar.
  • zdw 35 minutes ago
    Have you benched this for coding tasks, with a fallback to a larger local model, for example Qwen-3.6-27B?

    Or using it for sub-tasks, where a framework with a larger primary model dispatches simpler jobs ("summarize ...", etc.) to it?

  • cacio-e-pepe 6 hours ago
    > So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.

    Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)?

    Curious to learn more. Any more info on your approach (esp the mechanistic study)?

    • HenryNdubuaku 4 hours ago
      Correct, the study is verbose, we will compile into a neat shareable report and publish once we solve the pending caveats. Interesting username btw haha.
  • huflungdung 16 minutes ago
    [dead]