(Why) Was the LLM breakthrough useful for images, audio, etc.?

(Why) Was the LLM breakthrough useful for images, audio, etc.?

3 pointsby rogerrogerr5 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention
  • While some neural network architectures can be designed to better fit certain tasks they are at their core general-purpose learning algorithms which can approximate any target function.

    > I feel like I have a decent conceptual grasp of what LLMs are doing with written text

    Just think of it as input data. In theory it shouldn't matter what each token represents. They could be xbox controller buttons, image pixels, or text.

    The model with enough training data will map those inputs to an expected output.

  • VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.

    - https://huggingface.co/blog/vlms

    - https://en.wikipedia.org/wiki/Multimodal_learning

Explore Birbla archives