How to evaluate GM performance on your MU Online server
Build an objective process for evaluating Game Masters on your MU Online server, with support metrics, power-abuse indicators, and a monthly scorecard model for deciding promotions and terminations.
Keeping a team of Game Masters (GMs) motivated and trustworthy is one of the biggest challenges of running a private MU Online server: they're the ones deciding bans, resolving tickets, and holding access to commands that, if misused, can destroy community trust in a matter of hours. Without a struc
Keeping a team of Game Masters (GMs) motivated and trustworthy is one of the biggest challenges of running a private MU Online server: they're the ones deciding bans, resolving tickets, and holding access to commands that, if misused, can destroy community trust in a matter of hours. Without a structured evaluation process, decisions about promoting, warning, or firing a GM end up being emotional — based on a single complaint or personal sympathy. This tutorial presents an objective evaluation model, with measurable metrics, abuse-risk indicators, and a ready-to-use scorecard to apply monthly to your staff.
Why evaluate GMs in a structured way
A GM has access to commands that directly alter the game: teleport, item spawning, stat changes, account bans. Without formal evaluation, the server owner only discovers a problem once it has already turned into a public complaint on Discord or the forum. A structured process gets ahead of these problems: it catches drops in service quality, misuse of commands, and favoritism before the community notices. It also gives the GM a clear growth path — from Helper to full GM, from GM to Head-GM — based on visible criteria instead of the owner's favoritism.
Support metrics (volume and quality)
The first dimension of the evaluation is operational: is the GM resolving the tickets that come to them, and at what quality?
| Metric | How to measure | Suggested target |
|---|---|---|
| Average first-response time | Timestamp from ticket creation to the GM's first message | Under 10 minutes during peak hours |
| First-contact resolution rate | Tickets closed without reopening / total | Above 70% |
| Reopened tickets | How many tickets players complained about incomplete resolution | Below 10% |
| Volume per shift | Tickets handled / tickets available in the shift | Compare between GMs on the same shift, never across different shifts |
| Satisfaction score (if a post-ticket rating system exists) | Average 1-to-5 rating given by the player | Above 4.0 |
Pull this data from your own ticketing system (a Discord ticket bot, or a web panel) — most support bots already generate this kind of report automatically.
Power-abuse indicators
This is the most sensitive dimension, and the one that requires reliable server logs. Enable (or confirm you already have enabled) the GM command log in your emulator — every /summon, /additem, /ban, /setstats should record who executed it, when, and against which target.
| Indicator | What to check | Warning sign |
|---|---|---|
| Items generated for oneself | /additem commands targeting the GM's own account | Any occurrence outside documented testing |
| Bans without a recorded justification | Ban applied without an associated ticket or report | More than 1 case per month |
| Teleports outside a support context | /move//warp without an open ticket at that time | A recurring pattern, not an isolated case |
| Favoritism in disputes | Decisions consistently favoring the same group of players | Recurring complaints from several distinct players |
| Command usage outside the shift window | Command logged outside the declared shift window | Always investigate — could be a compromised account |
A single item generated "for testing" can be legitimate if documented in the staff channel before the action. The problem is the silent pattern, not the isolated event.
Evaluating communication and conduct
Beyond the numbers, qualitatively review a sample of the GM's conversations with players (with prior consent agreed on at hiring). Objective criteria:
- Uses respectful language even with an upset or hostile player.
- Explains the decision (why the ban, why the item won't be recreated) instead of just applying the action.
- Doesn't debate politics, religion, or personal matters in an official channel.
- Escalates to the Head-GM when a case is outside their scope, instead of improvising.
- Maintains the server's standard informal tone without offensive slang.
Monthly scorecard model
A simple scorecard, scored 0 to 5 per category with a different weight per dimension, works well for servers of any size:
| Category | Weight | Score (0-5) | Weighted score |
|---|---|---|---|
| Ticket volume and resolution | 25% | — | — |
| Support quality (communication) | 25% | — | — |
| Absence of abuse indicators | 30% | — | — |
| Shift/schedule compliance | 10% | — | — |
| Collaboration with the rest of the staff | 10% | — | — |
A final score below 2.5 in any month triggers a mandatory one-on-one conversation. Two consecutive scores below 2.5 in the "absence of abuse indicators" category justify suspending GM access, regardless of the other scores.
Feedback structure (monthly 1:1)
Calculating the score isn't enough — feedback needs to reach the GM constructively. A format that works: start with specific strengths (not generic ones — "quickly resolved the item-duplication case"), then improvement points with a concrete example (not "be more polite," but a screenshot of the problematic conversation), and end with an action plan for next month. Document the 1:1 in a private staff HR channel — this protects both the GM and management in case of a future dispute.
Criteria for promoting a Helper to GM
Many servers use a Helper → GM → Head-GM hierarchy. Define clear promotion criteria to avoid it feeling arbitrary:
| Criterion | Minimum requirement |
|---|---|
| Time in current role | 60 to 90 days |
| Average scorecard rating | Above 4.0 over the last 2 months |
| Zero confirmed abuse indicators | No occurrences during the period |
| Technical knowledge | Mastery of next-tier commands, tested in a staging environment |
| Recommendation from the direct Head-GM | Mandatory |
Criteria for warnings and termination
Just as promotion needs criteria, so does termination — this prevents the decision from looking like personal persecution and protects the server from disputes within the staff itself.
| Situation | Recommended action |
|---|---|
| First misconduct complaint, no proof of abuse | Verbal warning recorded in minutes |
| Second complaint in the same quarter | Formal written warning + temporary access suspension |
| Item generated for self/others confirmed in logs | Immediate termination, regardless of prior history |
| Leak of internal information (event plans, fixes) | Immediate termination |
| Performance drop with no sign of bad faith | 30-day improvement plan before any termination |
Tools to automate data collection
The more manual the collection, the less consistent the evaluation will be. Some practical options:
- Discord ticket bot (e.g., Ticket Tool, Hexadecimal) automatically generates response-time and volume reports.
- The emulator's own command log (MuEmu, IGCN) is usually stored in a SQL table — a simple query extracts commands by GM and by period.
- A shared spreadsheet (Google Sheets) with the scorecard, updated at month-end, keeps the history accessible to the whole leadership team.
- A private staff Discord channel with a webhook logging sensitive commands, for real-time auditing without needing to touch the database.
Handling conflicts between the evaluated GM and the evaluator
When the Head-GM has a personal relationship with the person being evaluated, the process loses credibility. Set a simple rule: no evaluator should evaluate a GM they have a personal relationship with outside the server (family, business partner in another project). In these cases, the server owner takes over the evaluation directly or delegates it to another Head-GM. Document this rule in the internal staff manual so it doesn't look improvised when the conflict actually arises.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Evaluation is always subjective and turns into an argument | Lack of objective metrics and logs | Enable the GM command log and adopt the weighted scorecard |
| GM feels targeted after a bad evaluation | Feedback given without concrete examples | Always attach a specific screenshot/log when pointing out a problem |
| Abuse is only discovered through a player report | No proactive log auditing | Do a light weekly review of sensitive commands, not just monthly |
| Promotions look like favoritism | Promotion criteria not documented | Publish the criteria in the internal manual before the next promotion |
| Ticket data doesn't match between GMs | Different measurement systems per shift | Standardize a single ticket bot/panel for all staff |
GM evaluation checklist
- GM command log active and accessible for auditing.
- Monthly scorecard with defined weights applied to all GMs.
- Monthly 1:1 scheduled with documented feedback.
- Promotion and termination criteria published in the staff manual.
- Conflict-of-interest rule between evaluator and evaluated defined.
- Proactive audit channel (not just reactive to reports) running.
- Evaluation history archived for reference in future disputes.
With the evaluation process running, the natural next step is to also review how the staff communicates with the server as a whole, including automated moderation channels — which connects directly to the overall admin structure described in the MU Online server creation tutorial.
Frequently asked questions
How often should I evaluate GMs?
Ideally, a light weekly review (ticket volume, response time) plus a full monthly review with the scorecard. Quarterly-only reviews let power abuse or quality drops go unnoticed for far too long.
Should I tell the GM they're being evaluated?
Yes, transparency about the process should be agreed on from the moment the GM joins the team. A hidden evaluation breeds distrust, and if the GM finds out later, it undermines staff authority with the rest of the team.
Which logs do I need to enable to evaluate correctly?
At minimum, the GM command log (item drops, teleports, ban/unban, stat changes) and the staff channel chat log. Without these two, any evaluation becomes guesswork and won't hold up against a favoritism complaint.
Is a GM with few resolved tickets always a problem?
Not necessarily — it may reflect a low-traffic shift. Always compare against the average number of tickets available in that shift, not against a fixed absolute number, or you'll unfairly penalize GMs on slow shifts.
What should I do when the evaluation reveals confirmed power abuse?
Suspend GM access immediately while you investigate, without waiting for the next evaluation cycle. Power abuse (giving items to yourself/friends, targeting a player) is the only category that requires action outside the normal schedule.