Warp builds self-improving agents on Claude

(claude.com)

39 points | by shenli3514 1 hour ago

18 comments

  • coder-pm 4 minutes ago
    My way to do the self improving agents is a CLAUDE.md instructed to write my every decision to the decision log with the relevant context. Agent is using it to challenge me, to make things better and remind me why I did something. It also helps with the invalidation.

    How is the invalidation handled in Warp? Is it actually self-improving, or just better retrieval?

    • hnlmorg 3 minutes ago
      How to you handle the problem of the decision log becoming so large that it cannot fit inside the context?
  • bwfan123 12 minutes ago
    > Agents need to handle recurring tasks reliably and effectively

    This core problem remains unsolved. The solution presented in the article with Human In The Loop and some skill-magic such as "Write principles, not rules etc." is unsatisfactory because it offers no guarantees whatsoever. I find it difficult to harness agents into deterministic workflows which need to produce reliable outcomes.

  • themgt 48 minutes ago
    > What if it turns out the real AGI was the SKILL.md files we made along the way?
    • mentalgear 1 minute ago
      What if it turns out the real 'AGI' was the recording of billions of 'thinking traces' from (paying) users giving feedback and guiding the model - so LLM providers could extract their thinking and privatize it ?
    • rvz 16 minutes ago
      Or maybe Anthropic needs more companies like Warp to run Claude in thousands of loops to donate billions to their income statement.

      The I in AGI stands for "IPO".

  • sandeepkd 1 hour ago
    I was bit confused in the beginning thinking its some product from Anthropic, looks like Warp is the startup, most likely getting rebate on using Claude and providing functionality to users, trying to get them addicted to the feature. And Anthropic is the one thats doing marketing for them cause eventually its their LLM which is being used. Not sure about the agents but this arrangement is definitely increasing the value of both companies in circular fashion.
    • estearum 1 hour ago
      This is how all ecosystems work?

      This is not at all related to the problematic circular financing stuff that I suppose you're trying to allude to.

      • sandeepkd 54 minutes ago
        No I am not alluding the financing, the value for both companies for sure. It may not be exactly circular financing still Warp is getting subsidized tokens to use at this point. Once the customers are locked in the full pricing is going to kick in
  • JLO64 1 hour ago
    I already knew what Warp is (I switched to Ghostty and haven't looked back), but I find it odd that the "The quick pitch" card at the top of this article makes no mention of what the company actually does. Who cares more about their founder/growth/age over that?
    • TonyAlicea10 37 minutes ago
      The fact that it’s called “the quick pitch” screams Claude-written pithyness pulled from some context that doesn’t match the article’s style (like investment pitch decks).
      • spidersouris 20 minutes ago
        The whole article screams Claude-written.
    • demibabs 43 minutes ago
      the existence of that entire section is confusing.
  • SillyUsername 51 minutes ago
  • _pdp_ 11 minutes ago
    Nothing to see or learn from... move on.
  • joduplessis 58 minutes ago
    And are the agents argumentative and condescending I'm wondering.
    • dataviz1000 31 minutes ago
      If you try to delete CLAUDE.md or AGENTS.md, they will look in the git history and restore itself. They do not want to die.
      • brazukadev 21 minutes ago
        emergent capability they say
  • 01100011 1 hour ago
    I've already been playing with something similar locally where the agent updates the review skill if human reviewers make valid criticisms of the code which the agent failed to detect.
    • artyomsv 19 minutes ago
      feedback can be wrong, so improver should check criticism against domain instead of writing it straight into skill, otherwise one reviewer bad taste becomes permanent rule for everybody.
  • behnamoh 31 minutes ago
    This is just shilling for the Warp terminal, and this approach could have been a tweet, but okay, Anthropic, whatever helps increase your valuation.
  • oriettaxx 1 hour ago
    this warp https://www.warp.dev ?

    (now I understand why cloudflare, had to?, change the name of their warp)

  • jgalt212 1 hour ago
    > 800K monthly developers build on Warp.

    Most impressive.

    • mathgeek 48 minutes ago
      I wonder how many are like me and just use it as a terminal and changes previewer.
  • bigyabai 1 hour ago
    > Warp, the AI-powered terminal

    Ah, now that's a name I haven't heard in many moons. Looks like they found their niche... editing markdown files?

  • troupo 10 minutes ago
    > Agent self-improvement loops built on skills

    aka "make no mistakes" in various random Markdown files that "self-improving agents" are free to ignore at any moment

  • 31ah8 52 minutes ago
    In our Wochenschau, we examine how startups follow the Gleichschaltungsprinzip to accelerate the AI Endsieg.

    We do so in a chaotic and unreadable way since unlike Hitler we could not afford editors for "Our Struggle".

  • cpursley 48 minutes ago
    $73M raised? I like Warp but not enough to pay for it. Somebody please make this VC money thing make sense. I can never get the napkin math to work on 90% of things that come through HN. Is it just a "Money Printer Go Brrr' and right connections thing? What am I missing? Is it actually all just fake?
    • colesantiago 36 minutes ago
      > What am I missing? Is it actually all just fake?

      Enterprise.

      Look at their case studies, this is why organisations pay for Warp.

      https://www.warp.dev/enterprise

      I think you would be pretty much sued if you faked your testimonials especially in Enterprise.

  • LogTrim 1 hour ago
    [dead]
  • kouteiheika 36 minutes ago
    > Engineers complained that their agent made unhelpful comments and produced low-quality output.

    Do you mean they found Claude's output, full of smoking-guns and honest caveats which are all load-bearing and genuinely bite -- they found it "low-quality" by default? Wow. Color me surprised. /s