mod_unique_id.html revision 0cff0af4031298cb16c17c038082e798af99cb9e
780a4529501556bf31f57e98b6a8ba389fcf735estoddard<html xmlns="http://www.w3.org/TR/xhtml1/strict"><head><!--
780a4529501556bf31f57e98b6a8ba389fcf735estoddardXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
780a4529501556bf31f57e98b6a8ba389fcf735estoddard This file is generated from xml source: DO NOT EDIT
780a4529501556bf31f57e98b6a8ba389fcf735estoddardXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
780a4529501556bf31f57e98b6a8ba389fcf735estoddard--><title>mod_unique_id - Apache HTTP Server</title><link href="/style/manual.css" type="text/css" rel="stylesheet"/></head><body><blockquote><div align="center"><img alt="[APACHE DOCUMENTATION]" src="/images/sub.gif"/><h3>Apache HTTP Server Version 2.0</h3></div><h1 align="center">Apache Module mod_unique_id</h1><table cellspacing="1" cellpadding="0" bgcolor="#cccccc"><tr><td><table bgcolor="#ffffff"><tr><td valign="top"><span class="help">Description:</span></td><td><description>Provides an environment variable with a unique
780a4529501556bf31f57e98b6a8ba389fcf735estoddardidentifier for each request</description></td></tr><tr><td><a href="module-dict.html#Status" class="help">Status:</a></td><td>Extension</td></tr><tr><td><a href="module-dict.html#ModuleIdentifier" class="help">Module&nbsp;Identifier:</a></td><td>unique_id_module</td></tr></table></td></tr></table><h2>Summary</h2><summary>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>This module provides a magic token for each request which is
780a4529501556bf31f57e98b6a8ba389fcf735estoddard guaranteed to be unique across "all" requests under very
780a4529501556bf31f57e98b6a8ba389fcf735estoddard specific conditions. The unique identifier is even unique
780a4529501556bf31f57e98b6a8ba389fcf735estoddard across multiple machines in a properly configured cluster of
780a4529501556bf31f57e98b6a8ba389fcf735estoddard machines. The environment variable <code>UNIQUE_ID</code> is
780a4529501556bf31f57e98b6a8ba389fcf735estoddard set to the identifier for each request. Unique identifiers are
780a4529501556bf31f57e98b6a8ba389fcf735estoddard useful for various reasons which are beyond the scope of this
780a4529501556bf31f57e98b6a8ba389fcf735estoddard document.</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard</summary><h2>Directives</h2><p>This module provides no directives.</p><h2>Theory</h2>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>First a brief recap of how the Apache server works on Unix
780a4529501556bf31f57e98b6a8ba389fcf735estoddard machines. This feature currently isn't supported on Windows NT.
780a4529501556bf31f57e98b6a8ba389fcf735estoddard On Unix machines, Apache creates several children, the children
780a4529501556bf31f57e98b6a8ba389fcf735estoddard process requests one at a time. Each child can serve multiple
780a4529501556bf31f57e98b6a8ba389fcf735estoddard requests in its lifetime. For the purpose of this discussion,
780a4529501556bf31f57e98b6a8ba389fcf735estoddard the children don't share any data with each other. We'll refer
780a4529501556bf31f57e98b6a8ba389fcf735estoddard to the children as httpd processes.</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>Your website has one or more machines under your
780a4529501556bf31f57e98b6a8ba389fcf735estoddard administrative control, together we'll call them a cluster of
780a4529501556bf31f57e98b6a8ba389fcf735estoddard machines. Each machine can possibly run multiple instances of
780a4529501556bf31f57e98b6a8ba389fcf735estoddard Apache. All of these collectively are considered "the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard universe", and with certain assumptions we'll show that in this
780a4529501556bf31f57e98b6a8ba389fcf735estoddard universe we can generate unique identifiers for each request,
780a4529501556bf31f57e98b6a8ba389fcf735estoddard without extensive communication between machines in the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard cluster.</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>The machines in your cluster should satisfy these
780a4529501556bf31f57e98b6a8ba389fcf735estoddard requirements. (Even if you have only one machine you should
780a4529501556bf31f57e98b6a8ba389fcf735estoddard synchronize its clock with NTP.)</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe <ul>
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe <li>The machines' times are synchronized via NTP or other
780a4529501556bf31f57e98b6a8ba389fcf735estoddard network time protocol.</li>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <li>The machines' hostnames all differ, such that the module
5f6e75bc39f8d95c7495ed17e585597cd6bd7fbatrawick can do a hostname lookup on the hostname and receive a
780a4529501556bf31f57e98b6a8ba389fcf735estoddard different IP address for each machine in the cluster.</li>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard </ul>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>As far as operating system assumptions go, we assume that
780a4529501556bf31f57e98b6a8ba389fcf735estoddard pids (process ids) fit in 32-bits. If the operating system uses
780a4529501556bf31f57e98b6a8ba389fcf735estoddard more than 32-bits for a pid, the fix is trivial but must be
780a4529501556bf31f57e98b6a8ba389fcf735estoddard performed in the code.</p>
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>Given those assumptions, at a single point in time we can
780a4529501556bf31f57e98b6a8ba389fcf735estoddard identify any httpd process on any machine in the cluster from
780a4529501556bf31f57e98b6a8ba389fcf735estoddard all other httpd processes. The machine's IP address and the pid
780a4529501556bf31f57e98b6a8ba389fcf735estoddard of the httpd process are sufficient to do this. So in order to
780a4529501556bf31f57e98b6a8ba389fcf735estoddard generate unique identifiers for requests we need only
780a4529501556bf31f57e98b6a8ba389fcf735estoddard distinguish between different points in time.</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>To distinguish time we will use a Unix timestamp (seconds
780a4529501556bf31f57e98b6a8ba389fcf735estoddard since January 1, 1970 UTC), and a 16-bit counter. The timestamp
780a4529501556bf31f57e98b6a8ba389fcf735estoddard has only one second granularity, so the counter is used to
780a4529501556bf31f57e98b6a8ba389fcf735estoddard represent up to 65536 values during a single second. The
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe quadruple <em>( ip_addr, pid, time_stamp, counter )</em> is
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe sufficient to enumerate 65536 requests per second per httpd
780a4529501556bf31f57e98b6a8ba389fcf735estoddard process. There are issues however with pid reuse over time, and
780a4529501556bf31f57e98b6a8ba389fcf735estoddard the counter is used to alleviate this issue.</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
db848a70422b56cf0f15a47c37e17ebe05e2ce04stoddard <p>When an httpd child is created, the counter is initialized
ec59f4e92f66631ee3266c0d416a01ac92bdf06cstoddard with ( current microseconds divided by 10 ) modulo 65536 (this
780a4529501556bf31f57e98b6a8ba389fcf735estoddard formula was chosen to eliminate some variance problems with the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard low order bits of the microsecond timers on some systems). When
780a4529501556bf31f57e98b6a8ba389fcf735estoddard a unique identifier is generated, the time stamp used is the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard time the request arrived at the web server. The counter is
780a4529501556bf31f57e98b6a8ba389fcf735estoddard incremented every time an identifier is generated (and allowed
780a4529501556bf31f57e98b6a8ba389fcf735estoddard to roll over).</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
58fcae4a8816f4a4d87b60a00346d095d085fce7wrowe <p>The kernel generates a pid for each process as it forks the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard process, and pids are allowed to roll over (they're 16-bits on
780a4529501556bf31f57e98b6a8ba389fcf735estoddard many Unixes, but newer systems have expanded to 32-bits). So
780a4529501556bf31f57e98b6a8ba389fcf735estoddard over time the same pid will be reused. However unless it is
780a4529501556bf31f57e98b6a8ba389fcf735estoddard reused within the same second, it does not destroy the
780a4529501556bf31f57e98b6a8ba389fcf735estoddard uniqueness of our quadruple. That is, we assume the system does
780a4529501556bf31f57e98b6a8ba389fcf735estoddard not spawn 65536 processes in a one second interval (it may even
780a4529501556bf31f57e98b6a8ba389fcf735estoddard be 32768 processes on some Unixes, but even this isn't likely
780a4529501556bf31f57e98b6a8ba389fcf735estoddard to happen).</p>
780a4529501556bf31f57e98b6a8ba389fcf735estoddard
780a4529501556bf31f57e98b6a8ba389fcf735estoddard <p>Suppose that time repeats itself for some reason. That is,
780a4529501556bf31f57e98b6a8ba389fcf735estoddard suppose that the system's clock is screwed up and it revisits a
780a4529501556bf31f57e98b6a8ba389fcf735estoddard past time (or it is too far forward, is reset correctly, and
780a4529501556bf31f57e98b6a8ba389fcf735estoddard then revisits the future time). In this case we can easily show
that we can get pid and time stamp reuse. The choice of
initializer for the counter is intended to help defeat this.
Note that we really want a random number to initialize the
counter, but there aren't any readily available numbers on most
systems (<em>i.e.</em>, you can't use rand() because you need
to seed the generator, and can't seed it with the time because
time, at least at one second resolution, has repeated itself).
This is not a perfect defense.</p>
<p>How good a defense is it? Suppose that one of your machines
serves at most 500 requests per second (which is a very
reasonable upper bound at this writing, because systems
generally do more than just shovel out static files). To do
that it will require a number of children which depends on how
many concurrent clients you have. But we'll be pessimistic and
suppose that a single child is able to serve 500 requests per
second. There are 1000 possible starting counter values such
that two sequences of 500 requests overlap. So there is a 1.5%
chance that if time (at one second resolution) repeats itself
this child will repeat a counter value, and uniqueness will be
broken. This was a very pessimistic example, and with real
world values it's even less likely to occur. If your system is
such that it's still likely to occur, then perhaps you should
make the counter 32 bits (by editing the code).</p>
<p>You may be concerned about the clock being "set back" during
summer daylight savings. However this isn't an issue because
the times used here are UTC, which "always" go forward. Note
that x86 based Unixes may need proper configuration for this to
be true -- they should be configured to assume that the
motherboard clock is on UTC and compensate appropriately. But
even still, if you're running NTP then your UTC time will be
correct very shortly after reboot.</p>
<p>The <code>UNIQUE_ID</code> environment variable is
constructed by encoding the 112-bit (32-bit IP address, 32 bit
pid, 32 bit time stamp, 16 bit counter) quadruple using the
alphabet <code>[A-Za-z0-9@-]</code> in a manner similar to MIME
base64 encoding, producing 19 characters. The MIME base64
alphabet is actually <code>[A-Za-z0-9+/]</code> however
<code>+</code> and <code>/</code> need to be specially encoded
in URLs, which makes them less desirable. All values are
encoded in network byte ordering so that the encoding is
comparable across architectures of different byte ordering. The
actual ordering of the encoding is: time stamp, IP address,
pid, counter. This ordering has a purpose, but it should be
emphasized that applications should not dissect the encoding.
Applications should treat the entire encoded
<code>UNIQUE_ID</code> as an opaque token, which can be
compared against other <code>UNIQUE_ID</code>s for equality
only.</p>
<p>The ordering was chosen such that it's possible to change
the encoding in the future without worrying about collision
with an existing database of <code>UNIQUE_ID</code>s. The new
encodings should also keep the time stamp as the first element,
and can otherwise use the same alphabet and bit length. Since
the time stamps are essentially an increasing sequence, it's
sufficient to have a <em>flag second</em> in which all machines
in the cluster stop serving and request, and stop using the old
encoding format. Afterwards they can resume requests and begin
issuing the new encodings.</p>
<p>This we believe is a relatively portable solution to this
problem. It can be extended to multithreaded systems like
Windows NT, and can grow with future needs. The identifiers
generated have essentially an infinite life-time because future
identifiers can be made longer as required. Essentially no
communication is required between machines in the cluster (only
NTP synchronization is required, which is low overhead), and no
communication between httpd processes is required (the
communication is implicit in the pid value assigned by the
kernel). In very specific situations the identifier can be
shortened, but more information needs to be assumed (for
example the 32-bit IP address is overkill for any site, but
there is no portable shorter replacement for it). </p>
<hr/><h3 align="center">Apache HTTP Server Version 2.0</h3><a href="./"><img alt="Index" src="/images/index.gif"/></a><a href="../"><img alt="Home" src="/images/home.gif"/></a></blockquote></body></html>