12 comments

  • simonw 1 hour ago
    I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

    For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

    Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

    If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

    (In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

    Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

    • 0xdeafbeef 47 minutes ago
      Better to have golden tests. Even switch from cpu to gpu can give you different tokens
      • slickytail 41 minutes ago
        Not really relevant in the case of embedding models. We might note that compressing a lossless PNG to 99% quality JPEG will also give you a different embedding. You don't care about bit for bit equality, just the distance in embedding space. And the same text embedded on CPU vs GPU will give very very similar embeddings.
  • Nautman 1 hour ago
    It's also very neat that this can be used for "Jev"-like tasks with text and image.

    https://developers.google.com/edge/mediapipe/solutions/decis...

    • Zambyte 15 minutes ago
      I got this running locally, and funnily enough the exact example they have for classifying "Cancel my flight and refund my credit card immediately." failed. It said the request does not involve a payment, charge, or refund, with a p(true) of 0.22 (where true means it is financial). Laya was significantly more accurate (in this case) and almost as fast.
    • helephants 34 minutes ago
      This is cool thanks!
  • flockonus 1 hour ago
    Hats off to google for offering OSS (or at least open weights + license) a model that would be probably pretty closed to what they would ship in their Android phones.
  • aabhay 1 hour ago
    Note that unlike prior on device embedding models, this seems to be trained with MRL, not MatFormers, meaning you don’t get to shrink the model weights alongside the lower dimensional embeddings, unfortunately. Likely there’s not good research for how to do MatFormers for multimodal yet?
  • dcl 1 hour ago
    Would be good to see how it compares to the embedding models from https://www.voyageai.com/ for text. I have used these a few times in the past and have found them superior to the Qwen models compared to here.
    • nostrebored 1 hour ago
      It's been awhile since I've been in the space, but Voyage was never a serious contender outside of super-niche business domains. I suspect that this compares favorably in 9X% of use cases
      • dcl 1 minute ago
        Interesting, I'll have to check this out.
  • minimaxir 6 hours ago
    Finally. I was getting annoyed that there's been an inflection point in how LLMs/agents work but there hasn't been a good moderate-size embeddings model, and this one is multimodal too! 270M for text only is great compared to older embedding models, and a total 440M for text + vision is also fair.

    I also may or may not have a tool for much faster local embedding creation that I calibrated for EmbeddingGemma but didn't want to release until a better embedding model came along.

    • alberto467 2 hours ago
      Not just vision with video, but also audio, it really seems amazing.

      I’m not sure how it can handle vicinity of pairs of embeddings with for example some words and the audio where they’re spoken or an image where the text is handled. Building local multimodal search with this would be amazing. I’ve explored this stuff with CLIP and it’s interesting how image (but also audio) embedding carries both the clean “text” content information but also the stylistic and visual/audio tone information, the two can even kind of be linearly separated.

      • treetalker 31 minutes ago
        > Building local multimodal search with this would be amazing.

        (lawyer here) — I’m curious: for what you’re describing, wouldn’t the machine need to be constantly running/updating the embeddings to take updated and new files into account? If so, how would that computation load compare to, say, Spotlight constantly updating its index?

  • Juvination 48 minutes ago
    So what are some use cases people have found for running these sized multimodals on their device? What is it accurate on, and what is the hallucination rate like?
    • minimaxir 43 minutes ago
      The main one is mapping images to text and visa versa, e.g. semantic search of images via text, where the images are encoded and the text question is encoded with the same model, then finding nearest neighbors.
    • treetalker 20 minutes ago
      I haven’t tried it yet, but, for law practice, I could imagine using it to search a case file for “undamaged roof before Hurricane Katrina” and “damaged roof after Hurricane Katrina” and being able to locate both deposition testimony and pertinent photographs in the body of evidence.
  • sourcecodeplz 59 minutes ago
    for text, benchmarks are identical to the first EmbeddingGemma.

    but you can use this new one and enable/disable what you don't need.

    can keep only text for ex.

  • djoldman 1 hour ago
    Parameter count split is interesting:

    740M total (270M text, 170M vision, 300M audio)

    • onlyrealcuzzo 1 hour ago
      Makes sense...

      Vision is just processing a still image (why it's by far the smallest). Text requires dealing with the entropy of human language. Audio is meaningless without time.

  • brokensegue 1 hour ago
    why isn't this being compared to siglip2 (also from google)? because that one isn't fully multimodal? or because it's a different org/team?
    • 392 16 minutes ago
      well EG2 doesn't have an encoder so can't use it for OCR, for one
  • nowittyusername 1 hour ago
    I'm considering adding this in my harness after some testing, this seems like a really nice embedding model!
  • sohamactive 1 hour ago
    rag transformations would be legendary