close

DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Installing GPT4All, an Open-Source Chatbot Application for Running LLMs

Installing GPT4All, an Open-Source Chatbot Application for Running LLMs

BERJAYA BERJAYA BERJAYA 10
Picked as gem Comments
9 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

Comments
3 min read
Every Greedy Metric Said the Model Was Improving. Then pass@64 Fell From 0.83 to 0.19

Every Greedy Metric Said the Model Was Improving. Then pass@64 Fell From 0.83 to 0.19

BERJAYA 1
Comments
4 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

Comments
4 min read
How to count the tokens in an LLM prompt (and why the number matters)

How to count the tokens in an LLM prompt (and why the number matters)

Comments
2 min read
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

Comments
6 min read
I built an AI API that turns any online store URL into a structured product catalog

I built an AI API that turns any online store URL into a structured product catalog

Comments
2 min read
I built an epistemic gate to stop LLM data poisoning during fine-tuning. Tested across 5 architectures, orchestrated on a 2006 Toshiba laptop for $0.

I built an epistemic gate to stop LLM data poisoning during fine-tuning. Tested across 5 architectures, orchestrated on a 2006 Toshiba laptop for $0.

Comments
2 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Comments
5 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

BERJAYA 1
Comments
4 min read
What an agent loop is (and isn't): state, action, stop

What an agent loop is (and isn't): state, action, stop

Comments
3 min read
Model routing with OpenRouter and DeepSeek to cut costs without losing quality

Model routing with OpenRouter and DeepSeek to cut costs without losing quality

Comments
3 min read
Reverse Engineering Undocumented Architectures: LLM-Driven Opcode Table Extraction vs. Legacy Tooling Constraints

Reverse Engineering Undocumented Architectures: LLM-Driven Opcode Table Extraction vs. Legacy Tooling Constraints

Comments
5 min read
Ollama's -cloud suffix isn't a label, it's a silent instruction. I bypassed it.

Ollama's -cloud suffix isn't a label, it's a silent instruction. I bypassed it.

BERJAYA 3
Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.