Sei sulla pagina 1di 2

PSYCHOMETRIKA--VOL. 40~ NO.

1
MARCH, 1975

GENERALIZED PROCRUSTES ANALYSIS

BY

J. C. GOWER

ROTHAMSTED EXPERIMENTAL STATION,


HARPENDEN~ HERTS.

Suppose piti) (i = 1, 2, • - • , m, j = 1, 2, • • •, n) give the locations of m n


points in p-dimensional space. Collectively these may be regarded a~s m
configurations, or scatings, each of n points in p-dimensions. The problem is
investigated of translating, rotating, reflecting and scaling the m configura-
tions to minimize the goodness-of-fit criterion ~2i= im E i~ ~" A~(pj(1)G~),where
G~ is the centroid of the m points Pitt) (i = 1, 2, • • • , m). The rotated posi-
tions of each configuration may be regarded as individual analyses with the
centroid configuration representing a consensus, and this relationship with
individual scaling analysis is discussed. A computational technique is given,
the results of which can be summarized in analysis of variance form. The
special case m = 2 corresponds to Classical Procrustes analysis but the choice
of criterion that fits each configuration to the common centroid configuration
avoids difficulties that arise when one set is fitted to the other, regarded
as fixed.

S u p p o s e X ~ ( i = 1, 2, . . . , m ) is a m a t r i x w i t h n rows a n d p~ c o l u m n s
whose j t h row gives t h e c o o r d i n a t e s of a p o i n t p ( 1 ) r e f e r r e d to p~ o r t h o g o n a l
axes. T y p i c a l p r a c t i c a l s i t u a t i o n s are w h e n each X i is a n o b s e r v e d d a t a -
m a t r i x or has b e e n o b t a i n e d as m u l t i d i m e n s i o n a l scales of n s t i m u l i or as
f a c t o r l o a d i n g s or scores etc. F i g u r e 1 i l l u s t r a t e s t h e g e o m e t r y for m = 3,
n = 4, Pl = P~ = P3 = 2 a n d w h e r e t h e t h r e e c o n f i g u r a t i o n s agree q u i t e well.
I n t h e following, i t is a s s u m e d t h a t t h e m p o i n t s P i " ) ( i = 1, 2, • • . , m )
all refer t o t h e s a m e j t h e n t i t y . F o r e x a m p l e , e a c h X~ m a y h a v e b e e n d e r i v e d
f r o m different sets of v a r i a b l e s o b s e r v e d for t h e s a m e s a m p l e s in e v e r y case
(j refers to samples) or e a c h . X , m a y h a v e b e e n o b t a i n e d as different seatings
of t h e s a m e o b s e r v a t i o n a l d a t a (j refers to o b s e r v a t i o n s ) or each X~ m a y h a v e
arisen f r o m one t y p e of scaling o f t h e s a m e s t i m u l i as p e r c e i v e d b y different
i n d i v i d u a l s (j refers to stimuli). I t is also a s s u m e d t h a t p~ = p, a c o n s t a n t .
T h i s simplifies t h e e x p o s i t i o n w i t h o u t loss of g e n e r a l i t y , b e c a u s e it is sufficient
t o choose p = M a x (p~) a n d a p p e n d zero c o l u m n s to e v e r y X~ t h a t i n i t i a l l y
h a s fewer t h a n p columns.
P r o b l e m s a r e c o m m o n w h e r e i t is d e s i r e d t o s t u d y t h e r e l a t i o n s h i p s
a m o n g s t t h e m sets a n d f r e q u e n t l y s o m e k i n d of c o m b i n e d a n a l y s i s is desirable.
E x a m p l e s are I n d i v i d u a l Scaling A n a l y s i s [Carroll & C h a n g , 1970] o r c o m -
p a r i n g factor l o a d i n g s d e r i v e d f r o m different studies. S c h S n e m a n n & C a r r o l l
33
34 PSYCHOMETRIKA

[1970] and Gower [1971] have discussed how two such matrices X1 and X2
can be fitted, allowing the motions of translation, rotation, reflection and
estimation of a homogeneous scaling factor, defining best fit as t h a t which
minimizes the least squares criterion m12 = ~ i = 1 ~ A2(P~ (1), pi(2)), where
5(.4, B) is the Euclidean distance between the pair of points A and B. This
problem has an analytical solution. For best fit, the centroids of X1 and X2
should be superimposed and the rotation matrix H such that X2H best
fits X~ is given as H = V'U where XI'X2 = U'FV, is the Eckart-Young [1935]
(or singular value) decomposition with U and V orthogonal and F diagonal.
With reflection F has no negative element and without reflection F has no
more than one negative element. The least-squares estimate of the scaling
factor obtained by fitting X2 to X~ is p = tr (X2HX~')/tr(X2X2') which is
clearly not the inverse of that obtained by fitting X~ to X2 . To overcome
this difficulty Sch6nemann & Carroll [1970] propose a symmetric scale measure;
an alternative solution occurs as a special case of the work described below.
Without translation and scaling this problem is known as Procrustes
rotation. In this paper the term Procrustes is extended to include all the
classical rigid-body motions and also the possibility of uniform scaling
(either stretching or shrinking). This extended terminology seems more in
the spirit of the Greek tale of the innkeeper Procrustes who stretched or
lopped off traveller's limbs so t h a t they would fit his bed.
Procrustes rotation provides one way of analyzing individually scaled
data. Given a pair of individual scales X~ and X. the first may be rotated
to best fit the second, giving a least-squares value m., ~ (say). With scaling
factors, in general, m., ~ m,. but without scaling factors m~. = m~ and it is
shown below t h a t elements of the (m X m) symmetric matrix of all such
comparisons, form a metric. This follows from considering three configura-
tions X1 , X2 and X3 and regarding all np values of each matrix (ordered by
columns, say) as being the coordinates of a single point, so that X I , X2 and X3
are represented by points Q~, Q2 and Q3 respectively, in a Euclidean np-space.
The distance between any two points Q. , Q~ of this representation is the
square root of the sum of squares (not necessarily minimal) measuring the
goodness of fit between X~ and Xo . If X1 and X2 have been rotated to best
fit with X3 , ml~ and m23 are merely the distances A(Q~ , Q3) and A(Q2 , Q~).
Because this representation is Euclidean, and hence metric,
A(Q~ , Q~) < rn,3 + m23 •
But A(Q~ , Q..) > m~ , because X~ and X2 cannot be rotated to a better fit
than that with sum of squares m122. Hence the metric inequality m~2 _<
m~3 ~ m23 must hold. Analysis of several sets of data suggests further that,
provided reflections are excluded, except perhaps for an initial reflection
of some of the X~ , the m~o form u matrix of Euclidean distances, but I have
not been able to prove this. The m X m matrix of all comparisons m a y be

Potrebbero piacerti anche