LibreTechni.ca
  • Communities
  • Create Post
  • Create Community
  • heart
    Support Lemmy
  • search
    Search
  • Login
  • Sign Up
cm0002@infosec.pub to AI - Artificial intelligence@programming.devEnglish · 30 days ago

Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era

joyemang33.github.io

external-link
message-square
1
fedilink
7
external-link

Humans Still Beat AI in the Long Horizon: Revisiting Test-Time Scaling in the Agent Era

joyemang33.github.io

cm0002@infosec.pub to AI - Artificial intelligence@programming.devEnglish · 30 days ago
message-square
1
fedilink
Agents can spend test-time compute by trying, observing, and revising. We derive an Elo reference for repeated sampling, then show that in a 2022 two-week coding marathon, current agents plateau within 24 hours while top humans keep improving.
alert-triangle
You must log in or register to comment.
  • Marija@programming.dev
    link
    fedilink
    English
    arrow-up
    1
    ·
    21 days ago

    AI is great at speed, not always direction.

AI - Artificial intelligence@programming.dev

Aii@programming.dev

Subscribe from Remote Instance

Create a post
You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: !Aii@programming.dev

AI related news and articles.

Rules:

  • No Videos.
  • No self promotion: Don’t post links to your articles.
Visibility: Public
globe

This community can be federated to other instances and be posted/commented in by their users.

  • 7 users / day
  • 170 users / week
  • 368 users / month
  • 1.03K users / 6 months
  • 1 local subscriber
  • 321 subscribers
  • 359 Posts
  • 368 Comments
  • Modlog
  • mods:
  • Vacant@programming.dev
  • cm0002@programming.dev
  • BE: 0.19.5
  • Modlog
  • Instances
  • Docs
  • Code
  • join-lemmy.org