// HACKER NEWS — CYBERSECURITY
What Would a Serious AI Product Look Like?
Every “AI” tool is missing critical features that you would need if
you wanted to do real work with them. What would those look like, if they
existed?
One of the issues that I have with the current generation of “AI” products is
that they do not appear to take their own premises seriously. I look at a
plethora of obsequious chatbots claiming to be serious tools for problem
solving, and I think, this is not what a problem-solving tool would look like.
Even before we get to the tremendous ethical problems with the frontier labs,
it is this impression of their composition as a product that makes me feel,
constantly, whenever I am interacting with them, that they are less a software
product than that they are a grift, a scam designed to make me feel like I am
interacting with a product that has capabilities that it simply does not, to
try to lull me into a false sense of security that I can trust it.
The frontier labs are of course the worst offenders, but every criticism here
applies just as much to Ollama, which (if anything, due to the obviously poorer
quality of the available models themselves) needs these features even more
than the frontier labs do.
Here, I will set down a few features that might convince me that an LLM-based
product, particularly one focused on research or software development, was
actually serious about helping me do useful things with it.
This is the biggest issue, and the major reason that I was inspired to write
this post.
It is a truth universally acknowledged, that AIs cannot reliably provide
information.
I could cite a ton of news articles and studies about this fact, but there is
no need. Every single chatbot admits this, up front, in a fine-print
disclaimer as a core part of their user interface. Gemini says “AI can make
mistakes, so double-check responses”, Claude says “Claude is AI and can make
mistakes. Please double-check responses.1” ChatGPT says “ChatGPT can make
mistakes. Check important info.”.
Every time I see that last one, I wonder how I’m supposed to know what “info”
is supposed to be “important”.
All of these warnings are all small, gray text, painfully obviously included as
legalese to push responsibility back onto the user rather than to help with
anything. This is a core limitation of all these products. Checking their
output is a part of the workflow for using them that: