‹ All posts

LangGraph and Langfuse, in plain words

My notes while learning two tools for AI apps. LangGraph decides what the app does next. Langfuse records what it actually did. You need both.

LangGraphLangfuseAI agents

These are my notes from learning these two tools. I made two small games from them — links at the end.

LangGraph: a map of steps#

A simple AI app is one prompt in, one answer out. An agent is more: it decides, uses a tool, checks the result, and maybe tries again. LangGraph lets you draw that as a map.

WordPlain meaning
NodeOne step. A small function.
StateThe notebook every step reads and writes.
EdgeAn arrow: what runs next.
Conditional edgeAn arrow that depends on what just happened.
ENDStop here.
python
graph.add_node("agent", agent)
graph.add_node("tools", tools)
graph.add_conditional_edges("agent", wants_tool, {"yes": "tools", "no": END})
graph.add_edge("tools", "agent")   # the tool's answer goes back to the agent

Loops need a way out#

Agent → tool → agent is a loop, and loops are what make agents work. But a loop that never ends burns money. LangGraph stops any run that takes too many steps (the recursion limit) with an error. Better: give the loop its own exit — “after 3 tries, hand it to a person”.

Pausing for a person#

Some steps should not happen without a human — like paying a big refund. LangGraph can pause right before that step, save everything (a checkpoint), and carry on from exactly there when someone says yes.

Langfuse: what really happened#

LangGraph decides what should happen. Langfuse records what did. Every run becomes a trace: a timeline of steps, with the time each took, the tokens and cost of every model call, the prompt version used, and any scores.

ComplaintWhere to look in the trace
“It’s slow”The longest bar on the timeline
“It’s wrong”What the search step returned
“It’s expensive”Input tokens — and what made them grow
“It changed yesterday”The prompt version on the model call
“It gave up”Steps marked ERROR

How I would test with both#

  • Keep a list of fixed test inputs, and the path each should take through the graph.
  • Run them in CI. Check the answer, and check the steps — a correct answer that took 9 tool calls is still a bug.
  • Send every run to Langfuse with its test name, so a failure comes with its full trace.
  • Add scores to traces (a judge label, user 👍/👎) and watch them after each prompt change.

Try it as a game#

Wire the agent yourself in the LangGraph game, then play trace detective in the Langfuse game. Each takes about five minutes.