this post was submitted on 05 Jan 2025
263 points (97.8% liked)

Technology

60291 readers
3346 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 2 years ago
MODERATORS
 

It's getting a bit ridiculous out here. I'm using DuckDuckGo but since it aggregates its search from other sources, it's also gotten bad recently. Is there a search out there that blocks domains that spam AI? Extra points if there's something like Ublock Origin that filters things based on a community-made list.

Edit: I'm aware of Kagi but it's pretty expensive and I'm not a fan that they, too, host their own AI tools.

you are viewing a single comment's thread
view the rest of the comments
[โ€“] rumba@lemmy.zip 7 points 2 days ago (1 children)

I tried doing some of this. I trained on a corpus of data I wanted it to read, with such a small amount of training data, I found it was overall too lossy. If I asked it a question about something that was in there and it responded there was a really good chance that it was in there. But there was a lot of not knowing something that was definitely in there. It wasn't completely useless but I wouldn't say that it was at the level of being truly helpful.

I worry that there's not enough verified data out there to set up for proper training.

I suspect such a model would have to be far more attuned to its data being smaller but trustworthy. Something like chatGPT for example requires a huge volume because it's weakly affected by any particular datum going in. It's designed to adapt to general conversation norms, rather than specific facts. If you could take a generalist like chatGPT and combine it with an expert model that's been told everything it's told has a huge weighting then that would probably be a big step forward.