BenchLab
or

By continuing you agree to BenchLab's Terms of Service.

Projects

Name Datasets Models Updated
LLM Benchmarking
128May 12
Vision Eval Suite
95May 10
RAG Evaluation
73May 8
Internal Tests
42May 6
1–4 of 4

LLM Benchmarking

Datasets

12

Models

8

Evaluations

48

Avg. Score

67.4

Recent Evaluations

  • May 12GPT-4o72.1
  • May 11Claude 3.564.8
  • May 10Llama 3 70B61.3

Leaderboard (Top 5)

  1. 1GPT-4o
    72.1
  2. 2Claude 3.5
    64.8
  3. 3Llama 3 70B
    61.3
  4. 4Mistral Large
    58.2
  5. 5Gemini 1.5 Pro
    55.7

Results

Dataset Model Score Price / 1k tok Updated

Evaluation Tasks

Drag cards between columns to update status.

To Do

3

In Progress

2

Done

4

File Manager

NameUpdated

Project created!

"Multimodal Eval" has been successfully created.

Settings / Profile

Alex Rivera

Alex Rivera

alex@domain.com