Saumitra Mazumder

On Jihad and Co.: A Game-Theoretic Perspective

I had some time over my vacation to read a treatise by Aisha Ahmad (incidentally, my former TA at U of T) titled: “Jihad & Co.: Black Markets and Islamist Power”.

The central argument, as I understand it, comes down to this passage:

[…] [T]wo specific economic logics that explain […] business Islamist cooperation: the need for higher trust and lower costs. First, in the long term, the business community adopts an Islamist identity to increase social trust and access to markets, which sets the stage for an Islamist takeover; second, in the short term, business elites make the strategic calculation to shift support to Islamist groups that provide them with better protection at a lower price, thus funding the Islamists to seize control of the state. These two rational, economic motivations lay the foundation of the business-Islamist alliance on which the Islamist proto-state can be built.

Reading it, I kept thinking back to a game theory course I took during my undergrad at TMU, and I couldn’t resist trying to put some of Ahmad’s argument into that language. Fair warning: it’s been a few years since I’ve done any serious game theory, so what follows is closer to “notes to myself, cleaned up enough to post” than a finished formal model. I’m not claiming to have captured everything in her book — just trying to see how much of the mechanism survives translation into a stylized model, and where it starts to strain.

A Game-Theoretic Analysis of Islamist Proto-State Formation

1. The Basic Idea

At its core, I think this is a story about security provision and coordination. Consider that in a weak or fragmented state, businesses can’t count on a central government to protect property, enforce contracts, or guarantee the movement of goods. This environment will require business to deal with whatever armed groups happen to control their patch of relevant territory, with each offering a different mix of protection and cost.

Per Ahmad, we question: under what conditions do a bunch of separate economic actors end up collectively shifting their support away from fragmented local security providers and toward one bigger organisation that can offer broader, cheaper security?

So what do we mean by collectively? Clearly, a single business might prefer one security provider over another. Then how much that preference is worth depends (perhaps) heavily on whether everyone else is making the same choice. A provider that can credibly operate across many territories only becomes valuable once enough businesses are actually relying on it.

In what follows, I’m going to assume both businesses and armed groups behave more or less rationally, that is, behavior on average responds to incentives closely enough that strategic patterns show up at the aggregate level.


2. Players

Take a simplified setup with two types of players:

Say provider $A$ is an ethnic or tribal armed faction whose support base is tied to a particular community, and provider $B$ is a faction capable of operating security across multiple communities — this is obviously a simplification of Ahmad’s actual cases, but it captures the contrast she’s drawing.

Each business has to pick who it deals with:

\[s_i \in \{A,B\}.\]

3. The Business’s Payoff

Let the payoff to business $i$ be

\[U_i = \pi_i - c_i - t_{s_i} - r_{s_i},\]

where:

So a rational business picks whichever arrangement maximizes expected return and hence a provider gets chosen when

\[t_{s_i}+r_{s_i}\]

is low enough relative to the alternative.

We can also consider that the cost and quality of a given provider isn’t fixed and make it depends on how many other businesses are also supporting it:

\[U_i(s_i,q)=\pi_i-c_i-t_{s_i}(q)-r_{s_i}(q),\]

with $q$ representing the share of businesses backing provider $B$.


4. Why Fragmentation Matters

Now,suppose provider $A$ is a patchwork of local factions tied to specific communities. A business operating across several territories then effectively deals with several different versions of $A$:

\[A_1,A_2,\ldots,A_k.\]

Its total security cost is then:

\[C_F(k)=\sum_{j=1}^{k}(t_j+r_j)+T(k),\]

where $T(k)$ represents any extra transaction costs that come from fragmentation itself, where in negotiating with $k$ separate groups (and perhaps conflicting groups) costs more than dealing with one. We’d expect

\[T'(k)>0.\]

Now, we can consider provider $B$ that can operate across the same territories. Here, cost depends on how big its network already is. Write

\[C_B(q,K)=t_B(q,K)+r_B(q,K)+T_B(q,K),\]

where $q$ is again the share of businesses on board and $K$ is the provider’s organizational capacity. The economic interpretation of this is that $B$ can spread its fixed costs over a bigger base as $q$ grows, and more capacity generally makes it better at actually delivering security, so

\[\frac{\partial C_B}{\partial q}<0, \qquad \frac{\partial C_B}{\partial K}<0.\]

$B$ becomes the better deal once

\[C_B(q,K)<C_F(k).\]

And hence a cross-cutting organisation can (perhaps) sell security to a much wider pool of businesses than one that’s boxed in by clan or ethnic lines.


5. Religious Identity

Per my interpretation of Ahamad’s argument, religious identity affects the cost of cooperation.

Let $\tau_{ij,t}$ be the trust between businesses $i$ and $j$ at time $t$. A shared identity, $I_{ij}$, might make cooperative behaviour more credible from the start, and a history of successful cooperation, $h_{ij,t}$, reinforces it further:

\[\tau_{ij,t+1} = (1-\delta)\tau_{ij,t} + \gamma I_{ij} + \eta h_{ij,t},\]

with $\gamma,\eta>0$ and $0<\delta<1$.

Identity makes cooperation more credible up front, and then repeated interactions do the rest of the work by building reputational capital. Cooperation holds up when the future value of the relationship outweighs the short-term gain from cheating,

\[\frac{g}{1-\delta} \geq d+\frac{p\ell}{1-\delta},\]

where $g$ is the ongoing gain from cooperating, $d$ is what you’d get by defecting once, and $p\ell$ is the expected reputational cost of getting caught.

Higher trust lowers transaction costs:

\[\frac{\partial C}{\partial \tau}<0.\]

Roughly, shared identity makes cooperation more credible, credible cooperation gets repeated, repetition builds reputational capital, and that capital lowers the cost of doing business together.

This also explains why an existing relationship with provider $A$ can be sticky even when $B$ looks better on paper. Businesses have often built up real reputational capital with $A$ over time and switching means giving that up.

Ahmad notes that shared identity doesn’t automatically produce trust, and there’s nothing intrinsically special about religious identity compared to other group identities that could do the same job.


6. The Security Provider’s Problem

The provider has its own strategic problem. Its revenue is

\[R=tN,\]

with $t$ the extraction rate per business and $N$ the number of businesses under its protection.

Raising $t$ brings in more per business while pushing some businesses away, so $\partial N/\partial t<0$, and the provider’s actual optimisation problem, once we account for the cost of maintaining its capacity $K$, is

\[\Pi_B(t,K;q) = tN(q,t,K)-C(K).\]

The provider picks $t$ and $K$ jointly:

\[(t^*,K^*) \in \arg\max_{t,K} \left\{ tN(q,t,K)-C(K) \right\}.\]

Where an interior solution exists, this gives the first-order conditions

\[N+t\frac{\partial N}{\partial t}=0\]

and

\[t\frac{\partial N}{\partial K}=C'(K).\]

That second condition defines that expanding capacity is possible since the extra revenue it brings in exceeds its marginal cost.


7. The Coordination Problem

Business 1 and Business 2 can each stick with incumbent provider $A$ or switch to $B$:

  Business 2: $A$ Business 2: $B$
Business 1: $A$ $(3,3)$ $(1,2)$
Business 1: $B$ $(2,1)$ $(4,4)$

This is a coordination game with two Nash equilibria, $(A,A)$ and $(B,B)$.

Here, $(B,B)$ is better for everyone, but a business that expects the other to stay with $A$ has every reason to stay with $A$ too, but moving off the one you’re already in requires everyone to move together.

To connect this discrete game back to something continuous, suppose both providers benefit from network effects:

\[U_i(A,q)=u_A+\beta(1-q),\] \[U_i(B,q)=u_B+\alpha q,\]

with $\alpha,\beta>0$ capturing how strong those network effects are for each side. The gap between the two options is

\[\Delta U(q) = U_i(B,q)-U_i(A,q) = (u_B-u_A-\beta)+(\alpha+\beta)q,\]

and a business switches to $B$ once $\Delta U(q)>0$. Solving for the threshold:

\[q^* = \frac{u_A+\beta-u_B}{\alpha+\beta}.\]

This is a coordination model where $B$ gets more attractive as more people adopt it, and $A$ gets less attractive as people leave. For there to be a real tipping to exist, we need

\[0<q^*<1,\]

which is equivalent to

\[u_B-\beta<u_A<u_B+\alpha.\]

If $u_B-\beta\geq u_A$, switching to $B$ is already dominant regardless of what anyone else does. If $u_A\geq u_B+\alpha$, $A$ stays attractive even if literally everyone else has left. It’s only in between that expectations about other people’s choices actually matter for the outcome.


8. A Continuous Coordination and Adoption Model

Extending this to a continuum of businesses, let $q\in[0,1]$ be the share supporting $B$. The payoff gap is now

\[\Delta U(q,t,K) = u_B(t,K)-u_A(t_A,K_A) + \alpha q - \beta(1-q),\]

where the first two terms are the “raw” attractiveness of each provider and the last two are the network effects layered on top. Each business’s best response is

\[BR_i(q,t,K) = \begin{cases} B,&\Delta U(q,t,K)>0,\\ A,&\Delta U(q,t,K)<0,\\ \{A,B\},&\Delta U(q,t,K)=0. \end{cases}\]

If businesses have heterogeneous switching thresholds, distributed according to $F$, aggregate adoption solves

\[q=F\!\left(\Delta U(q,t,K)\right).\]

And hence, the above is a fixed-point condition: whatever share adopts $B$ has to equal the share for whom adopting $B$ is actually optimal, given that everyone else is adopting at that same rate. In the homogeneous case this collapses back to the threshold from section 7,

\[q^{*} = \frac{u_A+\beta-u_B}{\alpha+\beta},\]

and as before, when $ 0<q^*<1 $, expectations below the threshold favor sticking with $A$ and expectations above it favor $B$ — consistent with the two-equilibrium story from the discrete game.


9. The Tipping-Point Dynamic

So assume there’s a threshold $ q^{} $ with $ 0<q^{} <1 $. Below it,

\[q<q^* \quad\Rightarrow\quad U_i(A,q)>U_i(B,q),\]

and businesses will tend to stay with $A$. Once expectations cross the threshold,

\[q>q^{*} \quad\Rightarrow\quad U_i(B,q)>U_i(A,q),\]

and support for $B$ starts feeding on itself.

From the abive, the tipping point is the moment where the combination of network effects and underlying costs makes switching the best response, given what everyone expects everyone else to do.

The entire feed foreword process is:

\[q\uparrow \Rightarrow N\uparrow \Rightarrow R\uparrow \Rightarrow K\uparrow \Rightarrow C_B\downarrow \Rightarrow U_i(B)\uparrow,\]

assuming the partial derivatives from earlier all point the way we assumed.


10. From Security Provider to Proto-State

The last step is going from “competitive security business” to something more like a proto-state.

Rather than bolting on a separate equation like $K=f(R)$, it’s more consistent to treat capacity as something the provider is already choosing optimally, per section 6:

\[t\frac{\partial N}{\partial K}=C'(K).\]

Higher expected demand raises the optimal capacity, but this relationship runs both ways rather than in one direction. If more capacity also lowers the effective cost of security ($\partial C_B/\partial K<0$), then adoption and capacity reinforce each other:

\[K\uparrow \Rightarrow C_B\downarrow \Rightarrow U_i(B)\uparrow \Rightarrow q\uparrow,\]

while at the same time $q\uparrow \Rightarrow N\uparrow \Rightarrow R\uparrow$, which gives the provider more room and more incentive to keep building capacity. So you get a loop: $q \to N \to R \to K^* \to C_B \to U_i(B) \to q$, and so on.


11. The Equilibrium

Pulling this together: an equilibrium is a tuple $(q^{},t^{},K^{*}) $ where each piece is optimal given the other two.

Businesses adopt according to

\[q^* = F\!\left( \Delta U(q^*,t^*,K^*) \right).\]

Provider $B$ picks extraction and capacity to solve

\[(t^*,K^*) \in \arg\max_{t,K} \left\{ tN(q^*,t,K)-C(K) \right\}.\]

Market size for $B$ follows directly, $N_B=Nq^$, giving revenue $R_B=t^Nq^$, while the incumbent $A$ is left with whatever’s left over, $N_A=N(1-q^)$.

Putting it in one place:

\[\boxed{ \begin{aligned} q^* &= F\!\left(\Delta U(q^*,t^*,K^*)\right),\\[4pt] (t^*,K^*) &\in \arg\max_{t,K} \left\{ tNq^*-C(K) \right\},\\[4pt] N_B^* &=Nq^*. \end{aligned} }\]

I don’t think this system necessarily has a unique closed-form solution. When $0<q^*_{\text{threshold}}<1$, you can have a low-adoption equilibrium where $B$ stays small and under-resourced right alongside a high-adoption equilibrium where a bigger market supports more capacity and lower costs.

That’s also, I think, the best explanation for why an inferior incumbent can hang on for a long time: switching requires everyone’s expectations to move together, and that’s a coordination problem, not just a pricing decision.


12. An Interpretation of the “Islamist Discount”

I think the “Islamist discount” idea can be restated as a cross-factional security advantage. If a tribal faction can only realistically sell security to $N_T$ businesses, while a cross-cutting organization reaches $N_I>N_T$, then fixed costs get spread more thinly for the bigger organization:

\[\text{average cost} = \frac{F}{N},\]

so if $N_I>N_T$, then $F/N_I<F/N_T$. That’s a real cost advantage, but it only exists if the organization can actually expand the market it credibly protects and taxes.

A broader market raises $q$, which raises $N_B$, which lets the provider build more capacity and lower its costs further. Fragmentation (section 4) and the tipping dynamic (sections 8–11) aren’t two separate stories; they’re the same cost advantage showing up at different scales.

Religious identity, in Ahmad’s account, is one plausible mechanism for that market expansion,where a cross-cutting identity that lets relationships extend past narrower clan or ethnic lines. I don’t think the model implies that Islamist identity is doing nothing economically real here. It’s just that its effect runs through more familiar channels, trust, network breadth, legitimacy, reputation, expectations — rather than being some sui generis force.


Some Thoughts

I think this model gives a plausible story for why some organizations succeed at this and others don’t. However, it’s not clear to me that an Islamist identity by itself guarantees (1) a broad security market, (2) credible cross-factional networks, or (3) enough organizational capacity to actually deliver on the promise. Ahmad’s own cases suggest these things co-occur more often than they’re guaranteed.

Some other provisos: existence, uniqueness, and stability of any of these equilibria depend entirely on the functional forms and parameter restrictions I’ve assumed, none of which I’ve derived from data and all of which I’ve derived from my intuition reading Ahmad.

When formal institutions collapse, people still have to solve the same recurring problems, trust, security, contract enforcement, coordination, and whoever solves them better tends to accumulate both economic and political support, as Ahmad’s cases illustrate. So maybe the right way to generalize her question —

Why do people support an Islamist organization?

— is something like:

When does an alternative security provider become credible and attractive enough that individual decisions to support it start reinforcing each other, and when does that process actually settle into a self-sustaining equilibrium of organizational capacity and political authority?

read more

On the coefficient of determination

On the coefficient of determination

I have often been asked by my students why the $R^2$ measure is not a measure of comparison between two models.

In general, as will be shown below, the $R^2$ value is simply a measure of how well the model that you’ve used fits the data that you have. In particular, when comparing nested models estimated by OLS on the same data, $R^2$ can never decrease when additional regressors are added.

Let’s start with a basic linear regression model:

\[Y = \mathbf{X}\beta + \varepsilon,\]

where $Y$ is an $n\times1$ vector, $\mathbf{X}$ is an $n\times k$ matrix of regressors, $\beta$ is a $k\times1$ vector of coefficients, and $\varepsilon$ is an $n\times1$ vector of disturbances.

We assume the usual conditions required for the OLS estimator to have its standard properties. In particular, if we are interested in interpreting the coefficients causally, we require appropriate exogeneity assumptions. Causality itself, however, is not one of the assumptions of the classical linear regression model.

When the relevant exogeneity and rank conditions hold, the least squares estimator for $\beta$, denoted $b$, is consistent. If $\mathbf{X}^{T}\mathbf{X}$ has an inverse, then $b$ is estimated via the least squares normal equations as follows:

\[b = (\mathbf{X}^{T}\mathbf{X})^{-1}\mathbf{X}^{T}Y.\]

Denote the vector of OLS residuals by

\[e = Y - \mathbf{X}b.\]

Substituting the expression for $b$, we have

\[\begin{aligned} e &=Y-\mathbf{X}(\mathbf{X}^{T}\mathbf{X})^{-1}\mathbf{X}^{T}Y\\ &=\left(\mathbf{I} -\mathbf{X}(\mathbf{X}^{T}\mathbf{X})^{-1}\mathbf{X}^{T}\right)Y. \end{aligned}\]

Define the residual-maker matrix

\[\mathbf{M}_{X} = \mathbf{I} -\mathbf{X}(\mathbf{X}^{T}\mathbf{X})^{-1}\mathbf{X}^{T}.\]

Then

\[e=\mathbf{M}_{X}Y.\]

Now, suppose we model the above with the addition of another variable:

\[Y = \mathbf{X}d + Zc + a,\]

where $Z$ is an $n\times1$ vector, $c$ is a scalar, and $a$ is the vector of residuals from the expanded model.

For a given value of $c$, the OLS estimator of the coefficient on $\mathbf{X}$ is

\[d = (\mathbf{X}^{T}\mathbf{X})^{-1} \mathbf{X}^{T}(Y-Zc).\]

Using the expression for $b$, this can be written as

\[d = b-(\mathbf{X}^{T}\mathbf{X})^{-1}\mathbf{X}^{T}Zc.\]

Hence,

\[\begin{aligned} a &=Y-\mathbf{X}d-Zc\\ &=\mathbf{M}_{X}(Y-Zc)\\ &=e-\mathbf{M}_{X}Zc. \end{aligned}\]

The Frisch–Waugh–Lovell theorem tells us that the coefficient $c$ in the expanded regression can be obtained by first removing the linear projection of $Z$ on $\mathbf{X}$, and the linear projection of $Y$ on $\mathbf{X}$, and then regressing the latter residuals on the former.

More importantly for our present purpose, the expanded model contains the original model as a special case. If we set $c=0$, then

\[Y=\mathbf{X}d+a,\]

and the original OLS solution is recovered by setting $d=b$.

The expanded model therefore minimizes the residual sum of squares over a parameter space that contains the parameter space of the original model. Consequently, the residual sum of squares from the expanded model cannot be larger than that from the original model:

\[a^{T}a \leq e^{T}e.\]

Finally, we arrive at $R^{2}$, also known as the coefficient of determination.

Suppose that the regression contains an intercept. Define $M^{0}$ as the $n\times n$ centering matrix

\[M^{0} = \mathbf{I} - \frac{1}{n}\mathbf{1}\mathbf{1}^{T},\]

where $\mathbf{1}$ is an $n\times1$ vector of ones.

The matrix $M^{0}$ transforms a vector into deviations from its sample mean. In particular,

\[Y^{T}M^{0}Y = \sum_{i=1}^{n}(Y_i-\bar Y)^2,\]

which is the total sum of squares.

For the regression

\[Y=\mathbf{X}b+e,\]

where $\mathbf{X}$ contains an intercept, the usual decomposition of the total sum of squares gives

\[Y^{T}M^{0}Y = b^{T}\mathbf{X}^{T}M^{0}\mathbf{X}b + e^{T}e.\]

The coefficient of determination is therefore

\[R^{2} = \frac{b^{T}\mathbf{X}^{T}M^{0}\mathbf{X}b} {Y^{T}M^{0}Y} = 1- \frac{e^{T}e} {Y^{T}M^{0}Y}.\]

Thus, $R^{2}$ measures the proportion of the sample variation in the dependent variable, relative to its sample mean, that is accounted for by the fitted regression.

Now consider the expanded model. Its residual sum of squares is $a^{T}a$, so its coefficient of determination is

\[R'^{2} = 1- \frac{a^{T}a} {Y^{T}M^{0}Y}.\]

Since

\[a^{T}a\leq e^{T}e,\]

it follows immediately that

\[R'^{2}\geq R^{2}.\]

That is, one can never decrease the ordinary $R^{2}$ by adding regressors to a nested OLS model, provided the same dependent variable and observations are used.

This is precisely why $R^{2}$ is not, by itself, a useful criterion for deciding whether one model is better than another. If we simply add more regressors to a model, $R^{2}$ will necessarily increase or remain unchanged, even when the additional regressors provide little substantive value.

This does not mean that adding regressors is necessarily harmful. It means only that ordinary $R^{2}$ does not impose any penalty for model complexity.

For this reason, when comparing models, one might instead consider adjusted $R^{2}$, information criteria such as AIC or BIC, statistical tests of restrictions, or out-of-sample predictive performance, depending on the purpose of the model.

In particular, if the objective is prediction rather than in-sample fit, an increase in $R^{2}$ need not imply an improvement in predictive performance.

read more

On ... and other ramblings

On … and other ramblings

This space will contain any short math or non-math thoughts that I have floating in my head.

read more