What the analyses look for¶
Clad generates a correct derivative for any code it can differentiate. How fast that derivative is depends on what its analyses can prove about the code, and each analysis proves it by recognising a construct: a pattern it knows how to exploit. Code that fits a construct gets the cheaper derivative. Code that does not still gets a correct one, built the conservative way.
This page lists the constructs, what each one buys, and what clad emits instead
when your code misses it. Each construct has a code, CLAD1001 and up, which
its reports print and which never changes; one section here per code.
Seeing what clad declined¶
Ask an analysis to report what it looked for and did not find:
clang -fplugin=clad.so -Xclang -plugin-arg-clad -Xclang -Rclad-analysis=loop ...
The name is one of tbr, activity, useful and loop, the same
names -fdisable-analysis= takes. Each report reads the same way: what
clad had to emit, what in the code forced it, and what to write instead.
gmm.cpp:53:3: remark: clad adds a counter to this loop and increments it every iteration
gmm.cpp:53:23: note: the bound is written elsewhere in the function
gmm.cpp:53:3: note: to avoid this, make it a counted loop (CLAD1001)
The caret of the first note is on what caused it, which is not always where the cost lands. Above, the cost is the loop; the cause is its bound.
To check whether an analysis is behind a wrong derivative, turn it off with
-fdisable-analysis=<name>, or turn off all of them with
-fdisable-analysis=all. Every analysis is optional by construction: with
all of them off clad still computes the same values, only more slowly. If a
derivative changes when you disable one, that is a bug in clad, and a good
bug report.
CLAD1001: a counted loop¶
An integer index stepped by one, from a start up to a bound, where nothing else in the function writes the start, the bound or the index:
for (int i = 0; i < n; i++) // counted, as long as the body leaves n and i alone
s += x[i] * y[i];
The reverse sweep of a loop has to run as many times as the loop did. For a counted loop clad works that number out from the bounds, so the forward sweep carries nothing. Otherwise it adds a counter, increments it every iteration, and for a nested loop pushes each count onto a tape, which allocates:
int k = n;
for (int i = 0; i < k; i++) { // not counted: the body writes k
s += x[i];
k = k - 1;
}
Rewriting the second loop to run to a bound nothing else touches is what makes the counter go away.
What a report says when this is missed:
the condition is not a comparison
the comparison is not < or <=
the left side of the comparison is not a variable
the index is not an integer
the header gives the index no start value
the increment does not step the index by one
the function can return before reaching this loop
the body can leave the loop early
the body assigns the index too
the start value is written elsewhere in the function
the bound is written elsewhere in the function
CLAD1002: a broadcast read¶
An element read on every iteration at an index the loop never moves, and read nowhere else in the loop:
for (int j = 0; j < n; j++)
s += x[i] * w[j]; // x[i] is a broadcast read of the j loop
The adjoint of a broadcast read is a sum over the loop. Clad keeps that sum in a register and adds it to memory once after the loop. Accumulated in place instead, it is a store to an address that does not change, and that one store is what stops the loop vectorising: LLVM reports it as a write to a loop invariant address:
for (int j = 0; j < n; j++)
s += x[0] * x[j]; // x is read at two indices, so neither moves
A read the loop’s own index moves is not a broadcast read at all, and clad says nothing about it.
What a report says when this is missed:
one element of this is not a number, and a register holds no more than one
this is read at more than one index in the loop
the loop writes this too
this is declared inside the loop, so each iteration gets a fresh one and its adjoint starts again
this is used somewhere other than as the base of a subscript, so something else may reach its adjoint
the index reads a variable the loop writes, so it is not the same element every iteration
CLAD1003: a bounded write¶
A write through a pointer parameter whose range clad can express in that function’s own parameters:
void subtract(int d, const double* x, const double* y, double* out) {
for (int i = 0; i < d; i++)
out[i] = x[i] - y[i]; // out is written over [0, d)
}
A call site that knows the range copies it once before the call and restores it after. Without it, clad falls back to watching every address the callee writes, which costs a record per element and a scan per restore.
The range has to be one range, over an index some counted loop steps, up to a bound the caller can work out: another parameter, or a constant.
What a report says when this is missed:
no counted loop steps the index of this write
the index is neither a constant nor a variable
two writes here cover ranges that are not one range
the bound is neither a parameter nor a constant, so a caller cannot work it out
this writes through a pointer that cannot be traced back to a parameter
this function has no body in this file
CLAD1004: a call clad can look inside¶
A direct call to a function with a body in this file.
The to-be-recorded analysis decides which values the reverse sweep will read back, so that the forward sweep saves only those. At a call it needs to know which arguments the callee reads and which it writes, and it learns that by analysing the callee. Where it cannot, every argument is saved.
Moving the definition into a header, or calling the function directly rather than through a pointer, is what lets the analysis see it.
What a report says when this is missed:
this function has no body in this file, so there is nothing to read to find out which arguments it reads
CLAD1005: a loop with a literal count¶
A counted loop whose start and bound are both literals:
for (int i = 0; i < 15; i++) // fifteen iterations, known while compiling
t = f(x, i); // t's fifteen values live in an array
What the body stores for the reverse sweep, one value per iteration, takes a slot in an array of that many at the loop’s index, and the reverse sweep reads it back by the same index, rather than pushing it onto a tape and popping it. The optimiser then sees an array where it saw calls into a tape. A loop clad cannot count, or one whose count is known only at run time, keeps the tape.
What a report says when this is missed:
the start or the bound is not a literal, so the count is not known until run time
Turning an analysis on or off
Each analysis is optional: it changes the code clad generates, never the
values that code computes. Turning one on asks clad to prove more and store
less; turning one off falls back to the conservative derivative.
-fenable-analysis=<name> and -fdisable-analysis=<name> take the names
below, and each analysis also has a switch of its own.
-enable-tbr/-disable-tbrTurns the tbr analysis on or off for the whole translation unit, unless an individual request specifies otherwise. It skips storing a value whose reverse sweep can recover it. Default: on.
-enable-va/-disable-vaTurns the activity analysis on or off for the whole translation unit, unless an individual request specifies otherwise. It skips the adjoint of a value no input can vary. Default: off.
-enable-ua/-disable-uaTurns the useful analysis on or off for the whole translation unit, unless an individual request specifies otherwise. It skips the adjoint of a value no output depends on. Default: off.
-enable-loop/-disable-loopTurns the loop analysis on or off for the whole translation unit, unless an individual request specifies otherwise. It proves what a function’s counted loops allow: a trip count the reverse sweep recomputes, and the range of a pointer parameter the body writes. Default: on.