Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I was pretty sure the "dynamic" in "dynamic programming" was used as a synonym for "awesome" and had nothing to do with system dynamics.
  • > An optimal policy has the property that whatever the initial state and initial decision are, the remaining decisions must constitute an optimal policy with regard to the state resulting from the first decision.

    So an optimal policy is an optimal policy. Got it.

  • side note: refreshing to read non ai articles
  • Thanks! What a great read.
  • >Proof: It all follows from applying the Banach fixed point theorem to the Bellman operator.

    I disagree. When people hate mathematicians, it's because of things like this.

    Why? Because it doesn't follow exclusively from the things listed. It also follows from the fact that c, T, A are bounded by the problem definition. Variable names in math confer no meaning, you can choose any variable name, hence you could have chosen an optimization problem that has no unique solution, simply out of spite.

    Now the counter to that is that the introduction covered the restrictions on c, A and T but why drop them in the proof?

  • Thanks for the refreshing reminder. This was one of my favourite topic at school but ended never used it in professional environment. However , I will definitely review the theory. Banach spaces and fix theorem were kind useful definition/tool but I always struggled to understand practical applications.

    Thanks for sharing the article

Explore Birbla archives