A music producer is the person accountable for turning an idea into a finished recording that other people can release, license, and rely on. The role covers writing decisions, session planning, performance direction, technical capture, mix balance, and the final file that leaves the studio. It is a creative role and a quality control role held by the same person, which is why the credit sits beside the artist rather than behind it.
Most of that work is measurable rather than intuitive. Sample rate, bit depth, recording headroom, frequency distribution, dynamic range, and integrated loudness are values that can be read on a meter and checked against a written specification. Taste decides which arrangement serves a song and which vocal take carries the intent of the lyric. Numbers decide whether the finished master is accepted by the platforms it is sent to, and a music producer answers to both.
Because of that, song production cannot be treated as a step added once the writing is done. Ownership terms, arrangement structure, recording standards, and delivery targets shape one another from the first session onward, and each of them closes options for the others. That is why a music producer belongs at the beginning of a project, not at the point where the mix is already being balanced and the structure can no longer be changed cheaply.
Music Producer and Music Production Services
DIMULTI Music works as a music producer in the full sense of the role, not as a booking service for studio hours. When we take on a song, we hold the whole chain, from the first writing session to the master that is uploaded to a distributor. One team carries the arrangement, the recording, the mixing and mastering, and the delivery, so no decision is left orphaned between vendors.
As a music producer, we treat production as a chain of constraints rather than a series of separate creative moments. Every decision made early narrows what remains possible later. A vocal captured in a room with the wrong reflections cannot be repaired with processing, an arrangement that leaves no frequency space cannot be rescued in the mix, and a mix built without a delivery target will change character once it is mastered.
What the client receives is a recording that behaves predictably outside the studio. It holds its balance on a phone speaker, in a car, and on a streaming platform that applies its own normalisation before anyone hears it. The same chain sits behind an artist single, a company anthem, and a short sound logo built for recall, because the standard a music producer applies does not change with the format.
Ownership, Credit, and Delivery Standards
A finished song carries more than one layer of rights, and Indonesian law separates them. Article 4 of Law Number 28 of 2014 defines copyright as moral rights and economic rights over the composition. Article 20 places a separate category beside it, Hak Terkait, which covers the performer, the phonogram producer, and the broadcaster. The song and the recording of it are two different properties.
Article 1 point 7 defines the phonogram producer as the party that first records the sound and carries responsibility for that recording, and Article 24 grants that party its own economic rights over reproduction, distribution, rental, and making the recording available. Article 63 sets the protection term for those rights at fifty years from the moment the phonogram is fixed. In practice that party is the record producer.
Transfers are formal. Article 16 states that economic rights may pass by written agreement, and the official elucidation of that article says the transfer must be made clearly and in writing, with or without a notarial deed. Moral rights stay attached to the author regardless. A verbal understanding about splits, reached in a studio at two in the morning, is not what the law is asking for.
Delivery standards bind the release just as firmly. A distributor expects a specific file format, accurate credits, and an ISRC, the twelve character identifier defined under ISO 3901 that combines a five character prefix issued by an ISRC agency, a two digit year, and a five digit designation code. DIMULTI settles scope, credits, ownership, and the final file list at the quotation stage, before the first session is booked.
Our Approach
1. Ownership Stated First
Master ownership and writer splits are fixed in writing before production begins, not after the song is finished. Percentages, credited roles, and the party holding the master, whether that is the artist, the label, or the music producer, are named in the quotation and signed before any session is scheduled.
This prevents the most common failure in a music project, which is a claim that surfaces only once the song starts to earn. Nothing has to be renegotiated under pressure, because the terms were agreed while the outcome was still unknown to everyone at the table. A music producer who postpones this is choosing to have the argument later.
For the client it means the song can be registered, licensed, and pitched without pausing to establish who is entitled to what. Publishing administration and sync opportunities can be pursued in the same week the master is delivered, rather than after another round of clarifying emails.
2. Arrangement Before Sound
Arrangement decisions are resolved before sound selection begins. Section order, harmonic movement, the entry point of the first chorus, and the density of each part are settled on a skeletal version, usually piano and voice, or guitar and voice, with nothing else in the way.
This prevents an expensive habit, which is using timbre to compensate for a structure that does not hold. A song that keeps attention with one instrument and one voice will keep it when fully produced. A song that does not will not be rescued by the sound palette layered over it.
The client hears the structural decision early, while changing it costs an afternoon rather than a full re-record. Approval lands at the point where revision is cheapest, and the remaining budget goes into performance and finish instead of repairing a shape that was settled too late. It is the cheapest hour a music producer spends on a song.
3. Revision Without a Ceiling, Scope With a Line
Revisions are not capped by number, because a music producer counting them is managing a budget rather than a song. Inside the agreed scope, a section is revised until it is right. What is fixed instead is the scope itself, written into the quotation as a specific list of deliverables, a song count, and an agreed arrangement direction.
This separates refinement from redefinition. Adjusting a vocal take, a balance, or a transition is refinement, and it stays inside the agreement. Changing the tempo, the key, the genre direction, or adding a second song is a different scope, and it is quoted as one before the work resumes.
The client never has to count revisions against a limit while trying to judge a mix. The only question that needs an answer is whether a change sits inside the scope written on the quotation, and that question has a written answer both sides agreed to in advance.
4. Release Ready by Definition
Finished means release ready, not studio ready. The master is measured against distribution specifications before it is called complete, including file format, sample rate, bit depth, true peak level, and integrated loudness measured with the ITU 1770 method, and it is checked on more than one playback system.
This prevents the most avoidable delay in a release, which is a master held or rejected by a distributor days before a scheduled date. It also prevents a master approved on the music producer’s own monitors from arriving quieter or thinner than the tracks placed around it in a playlist.
The client receives a package that can be handed to a distributor without an intermediate step. The same master serves the streaming release, while the accompanying versions cover later uses, such as an instrumental bed under a scene or a shortened cut for a campaign.
Our Solutions
1. Songwriting and Arrangement
Every project begins with the song rather than the sound of it. Melody, harmonic movement, lyric structure, and section order are the load bearing parts of a recording. They are also the parts a listener can still reproduce a week later, long after the production details have faded. A music producer settles them before a single instrument is chosen.
Arrangement is the plan for how attention moves through a song. It decides when an element enters, when one is taken away, and where the empty space sits that lets a chorus feel larger than the verse before it. Density is a variable a music producer controls on purpose, not a byproduct of how many parts happened to be recorded during the session.
The same chord progression carries a different meaning depending on how it is arranged. Instrumentation, register, rhythmic subdivision, and the number of simultaneous parts change how identical harmony is read. That is why a music producer builds more than one arrangement over a single progression before committing, and why the comparison is made out loud rather than described.
The opening of a song is not only a matter of taste. Hubert Leveille Gauvin examined 303 top ten singles released in the United States between 1986 and 2015 and measured how their openings changed across that period, in a study published in Musicae Scientiae in 2018. The direction of the change is consistent enough for a music producer to plan against.
Song Structure and Opening Length Overview
| Section | Working length | Function in the song | Decision it forces |
|---|---|---|---|
| Intro | 0 to 8 seconds | Establishes key, tempo, and texture | How long before a voice is heard |
| Verse 1 | 20 to 30 seconds | Sets the situation and the speaker | How much information the listener needs |
| Pre chorus | 8 to 15 seconds | Builds tension and signals arrival | Whether the chorus needs a run up at all |
| Chorus | 20 to 30 seconds | States the central idea and the title | Which line has to be memorable |
| Bridge | 15 to 25 seconds | Breaks the pattern before the final chorus | Whether the song has earned a contrast |
| Outro | 5 to 20 seconds | Releases tension or stops abruptly | How the track ends inside a playlist |
The lengths in this table are our working ranges, not rules. They exist so that a structural conversation with a client can happen in seconds and sections rather than in adjectives. A section that runs well outside its range is not wrong, but it has to earn the exception with a reason we can both name.
The Musicae Scientiae study measured one of these rows directly. Introductions that averaged more than twenty seconds in the mid nineteen eighties had fallen to roughly five seconds in the most recent part of the sample, a drop of about seventy eight percent. The time taken to reach the title fell by around eighteen percent, and average tempo rose by about eight percent.
Those figures describe commercial singles over three decades, not a target every song must meet. A ballad built for a wedding set and a track built for a playlist have different obligations. What the numbers do establish is that the length of an opening is a decision with consequences, and it is a music producer’s decision rather than something inherited from a demo.
How to Read the Section Table
- Read the second column as a range we work inside, not as a target duration to hit exactly.
- The third column is what the section does for the listener. If a section has no answer here, it is decoration and it can be cut.
- The fourth column is the question the section puts to the writer. It is the part that actually costs time in a session.
- The intro row is the one most affected by where a song will be heard. A track opening a live set can afford time a streaming single cannot.
- Pre chorus and bridge are optional rows. Many finished songs use neither, and adding one because the form seems to expect it is the most common structural mistake.
- Total length is the sum of the rows plus repeats. A four minute song and a two minute forty song are different arrangements, not the same arrangement trimmed.
Practical Application of Songwriting and Arrangement
The two examples below are the same song at two stages. The first is the skeletal version used for structural approval, with one instrument and one voice and no processing beyond level. Everything the listener responds to at this stage is structure, melody, and lyric, which is exactly what a music producer wants under discussion at this point.
The second is the same performance and the same structure with the full arrangement in place. The section boundaries have not moved. What has changed is density, register, and the amount of space around the vocal, which is the work of arrangement rather than the work of writing.
2. Recording and Vocal Production
Recording is the one stage that cannot be undone later. A mix can be rebalanced and a master can be redone, but a performance captured with the wrong microphone position, in a room with the wrong reflections, at a level that clipped the converter, stays that way. Everything downstream inherits whatever the recording chain wrote to disk, which is why a music producer settles that chain before the singer arrives.
Vocal production is a separate discipline from vocal recording, and it is where a music producer does most of the work. Recording is the capture. Production is the decision about which take carries the intent of the line, how the phrasing sits against the rhythm, where a breath is kept because it is human and where it is removed because it is noise, and how many layers a chorus needs before it stops sounding like one person.
Microphone distance is one of the few variables that changes tone before any processing is applied. Directional microphones raise low frequencies as the source moves closer, an effect of their pressure gradient design rather than a fault. Moving a singer closer produces weight and intimacy, moving them back produces clarity and room, and neither is a setting that a plugin can convincingly reproduce afterwards.
The technical side of capture is governed by published standards rather than preference. AES5, the Audio Engineering Society recommended practice on preferred sampling frequencies, names 48 kHz for professional origination, processing, and interchange. Its most recent edition was issued in 2018 and reaffirmed in 2023, and it is the reference our recording format follows.
Recording Signal Chain Overview
A performance travels to the recorded file through a fixed chain of stages: the room, the microphone, the preamp, and the converter. Each stage adds something that cannot be subtracted later, which is why the order and the settings are agreed before the singer arrives rather than adjusted while a take is running and the performer is waiting.
Two of the stages are irreversible in practice. The acoustic relationship between the source, the microphone, and the room is written into the file, and the analogue gain applied before conversion sets the noise floor and the headroom for everything that follows. The remaining stages can be revisited, so we keep them out of the recorded signal.
How to Read the Signal Chain
- The chain runs in one direction. Anything printed to the file at one stage is carried by every stage after it.
- The room is the first component in the chain, before the microphone. It is also the one most often ignored in a budget.
- Preamp gain sets both the noise floor and the headroom. Too little gain raises noise, too much removes the room for a loud phrase.
- Compression during capture is a decision, not a default. We record clean and shape afterwards unless there is a specific reason not to.
- Monitoring is separate from recording. What the singer hears in headphones can be processed freely, because none of it reaches the file.
- The recorded file is the deliverable of this stage. If it is right, mixing becomes a set of choices instead of a set of repairs.
Practical Application of Vocal Production
The pair below is one vocal take at two points in the process. The first is the raw file as recorded, with no editing, tuning, compression, or reverb. It is included because clients rarely hear this stage of a music producer’s work, and because the difference it shows is the difference vocal production actually makes.
The second is the same take after comping, timing and pitch correction held to the phrasing of the original, compression in two stages, subtractive equalisation, de essing, and reverb. Both files are matched in loudness so the comparison is about tone and control rather than about which one is louder.
Practice of Recording
We record at 48 kHz and 24 bit as standard, following AES5 for the sampling frequency. Twenty four bit quantisation provides a theoretical dynamic range of about 144 dB, far more than any room or microphone will deliver, and that margin is what allows conservative recording levels without audible penalty.
Levels are set so that peaks land around -12 dBFS and rarely above -6 dBFS. The headroom is a music producer’s insurance. A singer who commits to the last chorus can be six decibels louder than in rehearsal, and a chain set to run near full scale turns that commitment into a clipped file that no processing recovers.
Vocal distance normally sits between 15 and 25 centimetres with a pop filter, adjusted for the singer and the part rather than fixed for the session. Where a track needs weight we work closer and accept the low frequency rise. Where a track needs clarity we work further back and let the room contribute a small amount of natural space.
3. Mixing
Mixing is the allocation of a finite resource. Two instruments occupying the same frequency range at the same moment do not sum into something twice as impressive, they compete, and the listener hears the result as congestion. Most of what sounds like a lack of clarity in an amateur mix is simply two sources asking for the same space.
The work is therefore subtractive before it is additive. Deciding which element owns a range at a given moment, and removing the parts of the other elements that intrude on it, produces more separation than raising the element that is being buried. Volume is the last tool a music producer reaches for, not the first.
Dynamics are the second resource. Compression, automation, and arrangement all control how far a section can move between its quietest and loudest points, and how consistently a vocal sits above the music underneath it. A mix with no dynamic movement is fatiguing, and one with too much is unusable on a phone in a moving car, so a music producer sets that range against where the song will be heard.
The reason competing frequencies obscure one another is a documented property of hearing rather than a studio superstition. Harvey Fletcher described the pattern in Reviews of Modern Physics in 1940, setting out how the ear resolves sound into bands and how a tone within a band conceals a quieter tone that shares it.
Frequency Allocation Overview
| Range | What lives here | What competes | Decision taken |
|---|---|---|---|
| 20 to 60 Hz | Sub bass, kick fundamental | Room rumble, handling noise | High pass everything that does not belong |
| 60 to 120 Hz | Bass and kick weight | Kick against bass note | Alternate the range or duck one against the other |
| 120 to 300 Hz | Body of guitars, piano, vocal chest | Every layered instrument at once | Cut the buildup on supporting parts |
| 300 to 800 Hz | Snare body, lower vocal warmth | Boxiness from close microphones | Narrow cuts where the room resonated |
| 800 Hz to 3 kHz | Vocal intelligibility, lead lines | Lead synth or guitar against the voice | Reserve the range for the voice in vocal sections |
| 3 to 8 kHz | Presence, consonants, cymbal edge | Sibilance against hi hats | De ess the voice, soften the top of the kit |
| 8 to 16 kHz | Air and detail | Noise and hiss from every source | Add air on one element only |
The table is a plan for a specific song rather than a permanent map. A ballad with one voice and a piano has almost no competition to resolve, while a full band chorus with stacked vocals has competition in every row. What stays constant is the method a music producer follows, which is to name the owner of a range before adjusting anything.
Fletcher’s account of masking explains why this is worth the effort. Because the ear analyses sound in bands rather than at single frequencies, a louder sound conceals a quieter one sharing its band. Clearing the band is therefore more effective than raising the buried part, which only moves the competition to a higher level.
How to Read the Frequency Table
- Each row is a shared resource. Read the third column first, because it names the conflict the row exists to settle.
- Ownership can change between sections. A synth may own the midrange in an instrumental passage and give it back when the vocal returns.
- The decision column is usually a cut. Boosting a buried element leaves the conflict in place and raises the overall level.
- The lowest and highest rows are mostly cleanup. Very few sources need to contribute below 60 Hz or above 12 kHz.
- Sibilance sits in the same range as cymbal edge, which is why de essing and drum brightness are decided together.
- If a mix still feels crowded after the rows are settled, the problem is the arrangement, and it goes back a stage.
Practical Application of Mixing
The first file is a rough mix, which is a static balance with levels and panning set and nothing else. It is the version most projects mistake for a mix, and the one a music producer treats as a starting point. Nothing here is wrong, but every element is still asking for the same space, and the vocal survives only by being louder than everything around it.
The second is the finished mix of the same session. The frequency allocation in the table above has been applied, dynamics are controlled, and level automation follows the arrangement. The vocal is no longer louder in absolute terms, it is simply the only element occupying its range while it is singing.
4. Mastering and Delivery
Mastering is the last technical stage and the narrowest. It takes a finished mix and prepares it for the systems that will carry it, adjusting overall tonal balance, controlling peaks, setting final level, and producing the specific files each destination requires. It cannot repair a mix, and a master presented as a rescue is a mix a music producer will send back.
The decisive change in this stage arrived with loudness normalisation. Streaming platforms measure the integrated loudness of a track and adjust playback so that everything in a playlist arrives at a comparable level. A master pushed harder is therefore turned down on playback, and the only thing the extra limiting bought was a smaller difference between the quiet and loud parts.
That single mechanism reversed decades of practice. Under normalisation, the loud master and the moderate master play back at the same perceived level, and the moderate one keeps the transient detail the loud one destroyed. Judging them fairly means matching their loudness first, which is what a music producer does before a comparison is played.
The targets are published rather than inferred. Spotify states that it normalises to -14 dB LUFS using the ITU 1770 standard and asks that true peak stay below -1 dBTP, or below -2 dBTP if the delivered master is louder than the target. The Audio Engineering Society published its own recommendation for streaming in 2015, updated in 2021.
Loudness Target Overview
| Source | Stated figure | Applies to | Consequence for a master |
|---|---|---|---|
| Spotify, loudness normalization article | -14 LUFS integrated, ITU 1770 | Playback level of a delivered track | A louder master is attenuated on playback |
| Spotify, same article | -1 dBTP, or -2 dBTP above target | True peak of the delivered file | Headroom for lossy encoding artefacts |
| Spotify playback settings | -11, -14, and -19 LUFS | Listener selected playback level | The listening level is not fully in our hands |
| AES TD1004, 2015 | Not above -16, not below -20 LUFS | Target loudness a streaming service sets | Industry guidance sits below Spotify’s figure |
| AES TD1004, 2015 | Peaks not above -1 dBTP | Streams sent through lossy encoders | Same ceiling reached from a different direction |
The two sources in the table answer different questions. Spotify describes what its own service does to a delivered file. The Audio Engineering Society document recommends what target a streaming service should choose in the first place, and it names a window between -16 and -20 LUFS, measured with the true peak method defined in ITU-R BS.1770-3.
The AES document was published in 2015 as TD1004 and superseded in September 2021 by AESTD1008. Platform figures also change without notice. A music producer therefore checks the current published specification for each destination at the time a master is prepared, rather than working from a number remembered from a previous project.
How to Read a Loudness Target
- LUFS measures perceived loudness across a whole track. It is not the same as peak level, and two files with identical peaks can differ by many LUFS.
- Integrated means averaged over the entire track, so a quiet intro lowers the figure that a platform reads.
- dBTP is peak measured with oversampling, which catches peaks that appear after conversion and that a sample peak meter misses.
- A negative target is a ceiling for playback, not a score. Arriving under it costs nothing, because normalisation raises or leaves the track alone.
- Delivering above the target does not make a track louder to the listener. It makes it quieter after attenuation, with less dynamic range left.
- Compare two masters only after matching their loudness. Without matching, the louder file wins regardless of which one is better.
Practical Application of Mastering
The pair below is one mix in two masters, played at the same perceived loudness. The first is prepared for streaming delivery, with peak control applied conservatively and the transient detail of the drums left intact. This is the version a music producer delivers unless a client’s destination asks for something different.
The second is the same mix limited far harder, then attenuated so both files play at the same level. With the loudness advantage removed, what remains audible is what the limiting cost, which is transient definition on the drums and separation in the low midrange during the final chorus.
Practice of Delivery
The standard package is a WAV master at 24 bit and 48 kHz, an MP3 at 320 kbps for review and internal use, an instrumental version, and a version without the lead vocal for live use or for a guest feature. Stems are supplied on request, exported from a common start point so they align without adjustment.
Metadata travels with the files, because what a music producer delivers is a package rather than an audio file. Track title, artist, writer and performer credits, the year of the recording, and the ISRC are documented in a delivery sheet alongside the audio, so the distributor upload is a transcription task rather than a research task carried out under deadline.
Alternate durations are prepared at the same time when the release plan calls for them, because cutting a song to length is an arrangement decision for the music producer and it is cheaper while the session is still open. A shortened edit intended for a company profile or for a stage programme is planned into the session, not improvised from a finished master.
Our Experience as a Music Producer
DIMULTI Music has completed more than one hundred projects, and the majority of them ran through the full production chain described above rather than a single stage of it. That volume is the reason the process is written down. A method that is repeated a hundred times stops being a preference and becomes a standard the studio can be held to.
The range of collaborators runs from independent artists releasing a first single to teams operating at label scale, including producers who work internationally. For a music producer, the scale changes the schedule and the size of the approval chain. It does not change the recording format, the delivery list, or the ownership document.
Names we are able to mention include HITS, Yoda, Novia Bachmid, Lecrae, Dion, and Tovan. Each of those projects arrived with a different starting point, from a finished demo needing production to a brief with no melody attached yet, and each ended at the same place, which is a master with its rights and its files settled by the music producer who ran the session.
The work is not tied to one room or one time zone. Sessions run in the studio and remotely, with writing and revision handled asynchronously and tracking booked where the performer is. Releases from this desk have gone to streaming platforms, broadcast, corporate use, and event stages, including material built for seminar and event programmes.
What stays constant across all of it is unremarkable and deliberate. The brief is written before the session, the ownership is stated before the work, the files are checked against a specification before delivery, and the client is told what is happening while it is happening rather than after it has been decided.

Common Questions
Who owns the master when the project ends?
Whoever the signed quotation says owns it. In most artist and brand projects the client holds the master and DIMULTI retains the music producer credit. An institution commissioning a march or an anthem normally holds it outright. Under Article 16 of Law Number 28 of 2014 a transfer of economic rights is made in writing, so the document is drawn up before the first session.
Can a project be produced remotely?
Yes, and many are. Writing, arrangement approval, mixing, and mastering work asynchronously without loss, because each stage has a defined deliverable to review. Tracking is the stage that benefits from a room, so it is either booked here or recorded at a studio near the performer to an agreed format specification.
How long does one song take?
A music producer can quote the technical stages precisely and the approval stages not at all. Arrangement, tracking, mixing, and mastering occupy a known number of working days, and the schedule is quoted from that. What moves a timeline is waiting on a decision, which is why approval points are named in the quotation with the deliverable attached to each one.
What files arrive at the end?
A WAV master at 24 bit and 48 kHz, an MP3 at 320 kbps, an instrumental version, and a version without the lead vocal, together with a delivery sheet carrying the credits, the ISRC, and the loudness and peak figures the master was checked against. Anything beyond that list is agreed in the quotation.
Are stems included?
Stems are supplied on request and are stated in the quotation when they are part of the scope. They are exported from a common start point at the master format so they align without editing, which matters when they are later used for a live backing track, a remix, or a synchronisation cut.
What if a song needs a shorter advertising version?
Say so before the session rather than after the master. Shortened cuts are arrangement decisions a music producer makes, and they are cheap while the session is open and expensive once it is archived. When a campaign length is known in advance, the edit points are built into the arrangement and the cut is exported alongside the main master.
We Are Ready
A recording is no longer a document of a performance alone. It is an asset with rights attached to it, a technical specification it has to satisfy, and a set of derivative versions it will be asked for later. Treating production as the last creative step, rather than the stage that governs all three, is what leaves projects stalled at delivery.
Our answer is to combine the craft and the paperwork in one place. Writing, arrangement, recording, mixing, mastering, and the ownership document are handled by the same music producer and the same team, under one written scope, with the delivery specification agreed before the first note is recorded rather than discovered at the end of the process.
What the client ends up holding is a finished song they own on stated terms, in files that meet the specification of the places it is going, with the versions later requests will need already prepared. That is what engaging a music producer is supposed to deliver, and it is the standard every project through this studio is measured against.