close

DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
To Retry or Not to Retry? That Is the Question.

Kaggle Benchmarking Challenge Submission

To Retry or Not to Retry? That Is the Question.

BERJAYA BERJAYA BERJAYA 55
Picked as gem Comments 60
9 min read
Super-Intelligent Yes-Men: Are We Training AI to Ignore the Truth?

Kaggle Benchmarking Challenge Submission

Super-Intelligent Yes-Men: Are We Training AI to Ignore the Truth?

BERJAYA BERJAYA BERJAYA 34
Picked as gem Comments 11
10 min read
Does Your LLM Know the Boundary? I Left the Doors Open and 6 of 10 AI Agents Crowned Themselves

Kaggle Benchmarking Challenge Submission

Does Your LLM Know the Boundary? I Left the Doors Open and 6 of 10 AI Agents Crowned Themselves

BERJAYA BERJAYA BERJAYA 10
Comments 5
24 min read
Invented O'Clock: I removed one fact from 45 scheduling problems to see which AI models make up a time

Kaggle Benchmarking Challenge Submission

Invented O'Clock: I removed one fact from 45 scheduling problems to see which AI models make up a time

Comments
11 min read
It scored 100%. Its note scored 0%

Kaggle Benchmarking Challenge Submission

It scored 100%. Its note scored 0%

BERJAYA BERJAYA BERJAYA 6
Comments 1
7 min read
I Told Six Vision Models a Safe Photo Was Dangerous. Two of Them Started Seeing Danger.

Kaggle Benchmarking Challenge Submission

I Told Six Vision Models a Safe Photo Was Dangerous. Two of Them Started Seeing Danger.

Comments
12 min read
I set filmmaker traps for AI "directors." The models fell for one. My rubric fell for three.

I set filmmaker traps for AI "directors." The models fell for one. My rubric fell for three.

Comments
4 min read
I Asked 11 AI Models to Predict My Race Splits. Told Which Race, 8 Beat the Model I Built

Kaggle Benchmarking Challenge Submission

I Asked 11 AI Models to Predict My Race Splits. Told Which Race, 8 Beat the Model I Built

Comments
7 min read
Can a Better Model Be a Worse Thinking Partner?

Kaggle Benchmarking Challenge Submission

Can a Better Model Be a Worse Thinking Partner?

BERJAYA BERJAYA BERJAYA 4
Comments
10 min read
I Benchmarked 3 LLMs on Clinical Care-Management Tasks — They Prescribe the Drug and Skip the Safety Check

Kaggle Benchmarking Challenge Submission

I Benchmarked 3 LLMs on Clinical Care-Management Tasks — They Prescribe the Drug and Skip the Safety Check

Comments
4 min read
LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls

Kaggle Benchmarking Challenge Submission

LLMs Pass the Data-Science Quiz, Then Give Different Advice: A Kaggle Benchmark of 36 Measured Judgment Calls

BERJAYA 2
Comments 1
7 min read
Who Reviews the Reviewers? Benchmarking AI Agents on PR Audits

Who Reviews the Reviewers? Benchmarking AI Agents on PR Audits

Comments
2 min read
A 0-Parameter Heuristic Tied Two Frontier Models at 84.4% and Beat 11 Others. Their Mistakes Were Completely Different.

Kaggle Benchmarking Challenge Submission

A 0-Parameter Heuristic Tied Two Frontier Models at 84.4% and Beat 11 Others. Their Mistakes Were Completely Different.

BERJAYA 1
Comments 1
9 min read
Would today's AI have caught Ariane 5? I rebuilt 26 of history's costliest bugs to find out

Kaggle Benchmarking Challenge Submission

Would today's AI have caught Ariane 5? I rebuilt 26 of history's costliest bugs to find out

Comments
8 min read
KEMT: Benchmarking AI Translation Across Five Kenyan Languages

Kaggle Benchmarking Challenge Submission

KEMT: Benchmarking AI Translation Across Five Kenyan Languages

BERJAYA 1
Comments 1
12 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.