11:15 - 11:45 AM ET
Evaluation Is All You Need!
Jayeeta Putatunda
Large Language Models (LLMs) have transformed natural language processing (NLP), but their evaluation poses challenges due to the lack of standardized benchmarks for diverse tasks. The opaque, black-box nature of LLMs complicates understanding their decision-making processes and identifying biases. Effective evaluation metrics are crucial, especially as LLM architectures rapidly evolve, requiring adaptive methodologies. The AI community is coming together to address this, facilitate benchmark development, and provide tools for consistent model assessment across domains. We will also evaluate some of the OS evaluation metrics and walkthrough of code using a demo dataset.



















































