General Purpose Research

Shape update

Four new lessons and playgrounds to understand:

  1. Portability - how strong your steering is across vendors
  2. Context - how different reference documents affect output
  3. Tool calling - how to make sure your model uses the tools you expect it to
  4. LLM judge eval - if you hand assessments to an LLM, does it grade output like you expect it to?

Shape