Abstract
An observer of a process (xt) believes the process is
governed by Q whereas the true law is P. We bound the expected average
distance
between P(xt |x1, . . . , xt−1)
and Q(xt|x1, . . . , xt−1) for t = 1 .
. . n by
a function of the relative entropy between the marginals of P and Q on
the n first realizations. We apply this bound to the cost of learning in
sequential decision
problems and to the merging of Q to P.
Keywords
Bayesian Learning · Repeated Decision Problem · Value of Information ·
Entropy
Mathematics Subject Classification (2000)
62C10 62B10 91A26
Mathematical Programming, Springer Berlin /
Heidelberg, online first.